Staff Engineer, Inference Optimizations
$191,200–$239,000 year
RemoteSan Francisco, California, United States
Job Summary
Lead technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers to maximize throughput and minimize latency for large models. Engineer solutions for complex issues including attention layer optimizations, memory management, and advanced parallelization across multi-node GPU clusters. Proactively implement cutting-edge techniques such as kernel fusion for Transformer blocks and tuning expert gateway routers for MoE models. Act as a subject matter expert on NVIDIA and AMD GPU stacks, advising on hardware procurement and deploying state-of-the-art quantization techniques to double throughput. Partner with Product Management to translate theoretical hardware limits into shippable features while maintaining a strong presence in open-source AI communities.
Required Qualifications
- 5+ years of experience in high-performance computing or AI infrastructure
- proven track record of solving compute utilization and memory bandwidth bottlenecks
- Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape
- Deep familiarity with the specific quirks and architectural requirements of major model families
- Hands-on experience with attention-layer optimizations
- experience with parallelization strategies across distributed GPU environments
- Comprehensive understanding of NVIDIA and AMD GPU architectures
- Comprehensive understanding of NVIDIA and AMD software ecosystems (CUDA, ROCm, etc.)
- Extensive experience integrating, building with, and contributing to open-source software projects
- Excellent system design skills, particularly related to low-level GPU programming
- Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores)
- Expert-level Triton or CUDA
- contributed to the Triton compiler
- wrote custom CUDA kernels for a major LLM
- Experience acting as a technical lead
- driving design and delivery through cross-functional alignment
- expert-level delegation
Desired Qualifications
- experience with AMD's AITER library for AMD MI355X
- identify and tune AITER's CK (composable kernel) or ASK (assembly) to optimize FP8 / BF16
- Identify kernel fusion opportunities for GLM-5 kernels for different layers of the Transformer block (FlashAttention, RMS Norm)
- Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc
- Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.