GPU Software Engineer (CUDA)
$100,000–$175,000 year
RemoteUnited States
Job Summary
Design and implement high-performance CUDA kernels for compute-intensive workloads across AI and HPC use cases. Profile and optimize GPU code using Nsight Systems, Nsight Compute, and CUDA profilers to tune memory access patterns, occupancy, and shared memory utilization. Develop highly optimized libraries for linear algebra and ML primitives, while implementing custom operators in PyTorch, JAX, or Triton. Collaborate with cross-functional partners to translate requirements into engineered solutions, conduct code reviews, and mentor junior engineers. Evaluate new GPU architectures and feature sets to advise on adoption strategies. This role requires 6+ years of GPU programming experience and is open to U.S. citizens, Green Card holders, EAD holders, and H-1B transfer candidates. Compensation ranges from $100,000 to $175,000 annually for this remote position at Bright Vision Technologies.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field
- Six or more years of experience in GPU programming and performance engineering
- Deep expertise in CUDA C/C++ and GPU programming models
- Strong understanding of modern GPU architectures, memory hierarchies, and execution models
- Hands-on experience profiling and optimizing GPU workloads in production
- Familiarity with NCCL, MPI, and high-performance interconnect technologies
- Experience integrating custom kernels into ML frameworks
- Strong C++ skills and familiarity with modern systems programming practices
- Solid grounding in linear algebra and numerical methods
- Strong communication and collaboration skills with research and engineering teams
- 6+ years
- U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates
Desired Qualifications
- Experience with Triton, CUTLASS, or other GPU kernel authoring frameworks
- Familiarity with TensorRT, FasterTransformer, or vLLM internals
- Exposure to compiler infrastructure such as LLVM or MLIR
- Open-source contributions to GPU or ML performance libraries
- Experience with large-scale distributed training infrastructure
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.