Staff Engineer, Inference Optimizations
$191,200–$239,000 year
RemoteBoston, Massachusetts, United States
Job Summary
Lead technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers to maximize throughput and minimize latency. Engineer solutions for complex issues including attention layer optimizations, memory management, and advanced parallelization across multi-node GPU clusters. Proactively implement cutting-edge techniques such as kernel fusion for Transformer blocks and expert gateway router tuning for MoE models. Act as a subject matter expert on NVIDIA and AMD GPU stacks, developing state-of-the-art quantization techniques to double throughput. Drive design and delivery through cross-functional alignment, partnering with Product Management to translate hardware limits into shippable features while maintaining a strong presence in open-source AI communities.
Required Qualifications
- 5+ years of experience in high-performance computing or AI infrastructure
- proven track record of solving compute utilization and memory bandwidth bottlenecks
- Deep familiarity with the Gen AI (LLM, VLM, LMM) landscape
- specific quirks and architectural requirements of major model families
- Hands-on experience with attention-layer optimizations
- parallelization strategies across distributed GPU environments
- Comprehensive understanding of NVIDIA and AMD GPU architectures
- respective software ecosystems (CUDA, ROCm, etc.)
- Extensive experience integrating, building with, and contributing to open-source software projects
- Excellent system design skills
- low-level GPU programming - optimization, memory access patterns, and parallel execution
- Experience acting as a technical lead
- driving design and delivery through cross-functional alignment
- expert-level delegation
- Deep understanding of GPU architectures (SMs, Warp scheduling, Tensor Cores)
- Expert-level Triton
- wrote custom CUDA kernels for a major LLM
Desired Qualifications
- growth mindset
- naturally like to think big and bold
- energized by the fast-paced environment of a true industry disruptor
- Improving batch size performance using AMD's AITER library for AMD MI355X
- identify and tune AITER's CK (composable kernel) or ASK (assembly) to optimize FP8 / BF16
- Identify kernel fusion opportunities for GLM-5 kernels for different layers of the Transformer block (FlashAttention, RMS Norm)
- Tune expert gateway router kernels for MoE models like Qwen3-235B, DeepSeek V3, GLM-5 etc
- Develop and deploy state-of-the-art quantization techniques (FP8, INT8, and experimental FP4) to double throughput without losing accuracy
- Lead the technical strategy for benchmarking and performance optimizations at the inference engine and GPU kernel layers
- Ensure our infrastructure extracts maximum value from every TFLOP
- Engineer solutions for complex performance issues
- including attention layer optimizations, memory and precision management, and advanced parallelization across multi-node GPU clusters
- Proactively implement cutting-edge optimization techniques to keep DigitalOcean at the forefront of the Gen AI landscape
- Act as the subject matter expert on modern GPU families (NVIDIA/AMD) and their software stacks (CUDA, ROCm, TensorRT, OpenAI Triton)
- advising on hardware procurement and software integration
- Lead by example through high-quality code and design reviews
- elevating the technical bar for the team without the administrative overhead of direct management
- Partner with Product Management and TPMs to translate 'theoretical hardware limits' into 'shippable product features'
- ensuring our platform is both powerful and developer-friendly
- Maintain a strong presence in the GPU infrastructure and model performance optimization communities
- contributing to and integrating the best of open-source AI
- Shark who thinks big, bold, and scrappy
- like an owner with a bias for action
- a powerful sense of responsibility for customers, products, employees, and decisions
- We provide employees with reimbursement for relevant conferences, training, and education
- All employees have access to LinkedIn Learning's 10,000+ courses to support their continued growth and development
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.