Staff Engineer, Inference Optimizations
$191,200–$239,000 year
RemoteAustin, Texas, United States
Job Summary
Lead technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers to maximize throughput and minimize latency. Engineer solutions for complex issues including attention layer optimizations, memory management, and advanced parallelization across multi-node GPU clusters. Proactively implement cutting-edge techniques such as kernel fusion for Transformer blocks and expert gateway router tuning for MoE models like Qwen3 and DeepSeek. Develop and deploy state-of-the-art quantization techniques to double throughput while maintaining accuracy. Act as a subject matter expert on NVIDIA and AMD GPU stacks, advising on hardware procurement and software integration. Guide the technical roadmap for the high-performance inference fleet through design reviews and cross-functional collaboration with Product Management.
Required Qualifications
- 5+ years of experience in high-performance computing or AI infrastructure
- proven track record of solving compute utilization and memory bandwidth bottlenecks
- Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape
- Deep familiarity with the specific quirks and architectural requirements of major model families
- Hands-on experience with attention-layer optimizations
- experience with parallelization strategies across distributed GPU environments
- Comprehensive understanding of NVIDIA and AMD GPU architectures
- Comprehensive understanding of NVIDIA and AMD software ecosystems (CUDA, ROCm, etc.)
- Extensive experience integrating, building with, and contributing to open-source software projects
- Excellent system design skills, particularly related to low-level GPU programming
- Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores)
- Expert-level Triton or CUDA
- Experience acting as a technical lead
- driving design and delivery through cross-functional alignment
- expert-level delegation
- *This is a remote role
Desired Qualifications
- If you've contributed to the Triton compiler
- wrote custom CUDA kernels for a major LLM
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.