Software Engineer, CUDA Deep Learning Systems
$124,000–$195,500 year
RemoteAustin, Texas, United States or Santa Clara, California, United States
Job Summary
Explore, research, and prototype novel systems optimizations for advanced deep learning models at the intersection of high-level frameworks and low-level CUDA through modeling, simulation, and silicon prototyping. Architect and optimize distributed computing systems scaling from single nodes to massive cluster environments. Design, implement, and optimize custom high-performance CUDA kernels tailored to emerging neural network architectures. Analyze hardware-software interactions to resolve performance bottlenecks in training and inference pipelines. Collaborate with AI researchers, HW and SW architects, and compiler experts to co-design systems improving accelerator compute utilization and network efficiency. Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms. Write clean, maintainable code ensuring prototypes transition smoothly into open-source releases or commercial products.
Required Qualifications
- BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience)
- 2+ years of relevant industry experience or equivalent academic experience after degree achievement
- Strong proficiency in C++ and Python programming
- Solid background in the fundamentals of Deep Learning with a focus on transformers
- Strong understanding of distributed computing principles, multi-node scaling, and the unique performance challenges of cluster-scale execution
- Proven experience in systems programming, computer architecture, and low-level systems performance optimization
- Familiarity with deep learning accelerator architectures such as the GPU and hands-on experience with CUDA programming, kernel optimization, and workload profiling
- Experience profiling and optimizing generative AI models, including but not limited to, pioneering large language models
- Research background in machine learning systems or adjacent fields and experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models
- A track-record of initiative and willingness to deep-dive on problems across the stack
Desired Qualifications
- Deep expertise in performance internals and execution graphs of major deep learning training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron)
- Hands-on experience with communication libraries (e.g., NCCL, MPI, UCX) and distributed machine learning techniques (e.g., pipeline, tensor, expert parallelism)
- Knowledge of numerical methods and low-precision arithmetic (e.g., NVFP4, MXFP4, FP8, INT8) and their impact on deep learning accuracy and performance
- Background in deep learning compilers and ML systems, including graph-level and codegen tools (e.g., Triton, XLA, torch.compile) and highly parallel/RL-style simulation environments
- Experience designing and implementing agentic AI systems applied to complex systems and infrastructure problems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.