Staff Engineer, Inference Optimizations
$191,200–$239,000 year
RemoteDenver, Colorado, United States
Job Summary
Lead technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers, ensuring infrastructure extracts maximum value from every TFLOP. Engineer solutions for complex issues including attention layer optimizations, memory management, and advanced parallelization across multi-node GPU clusters. Proactively implement cutting-edge techniques such as tuning AMD AITER libraries for FP8/BF16, identifying kernel fusion opportunities for Transformer blocks, and optimizing expert gateway routers for MoE models. Act as a subject matter expert on NVIDIA and AMD GPU stacks, developing state-of-the-art quantization techniques to double throughput. Lead by example through high-quality code reviews and partner with Product Management to translate hardware limits into shippable features. Maintain a strong presence in open-source AI communities while guiding the technical roadmap for the high-performance inference fleet.
Required Qualifications
- 5+ years of experience in high-performance computing or AI infrastructure
- proven track record of solving compute utilization and memory bandwidth bottlenecks
- Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape
- Deep familiarity with the specific quirks and architectural requirements of major model families
- Hands-on experience with attention-layer optimizations
- experience with parallelization strategies across distributed GPU environments
- Comprehensive understanding of NVIDIA and AMD GPU architectures
- Comprehensive understanding of NVIDIA and AMD software ecosystems (CUDA, ROCm, etc.)
- Extensive experience integrating, building with, and contributing to open-source software projects
- Excellent system design skills, particularly related to low-level GPU programming
- Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores)
- Expert-level Triton or CUDA
- Experience acting as a technical lead
- driving design and delivery through cross-functional alignment
- expert-level delegation
- *This is a remote role
Desired Qualifications
- If you've contributed to the Triton compiler
- wrote custom CUDA kernels for a major LLM
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.