Senior Deep Learning Algorithm Engineer
$152,000–$241,500 year
RemoteCalifornia, United States or Santa Clara, California, United States
Job Summary
Design and maintain Dynamo integrations for open-source frameworks vLLM, SGLang, and TensorRT-LLM. Partner with communities to land measurable gains in latency, throughput, reliability, and efficiency. Find and remove bottlenecks across runtimes, kernels, networking, and orchestration while developing optimizations for scheduling, disaggregation, KV caching, and autoscaling. Showcase NVIDIA token/watt leadership by pushing the Pareto frontier on public and private benchmarks. Collaborate across research, software, systems, and hardware teams to make AI inference faster and easier to deploy.
Required Qualifications
- BS, MS, PhD in Computer Science, Electrical Engineering, Computer Engineering, or a related field (or equivalent experience)
- 3+ years building, profiling, and debugging performance-critical distributed or ML systems
- Strong programming skills in Python and/or Rust, C++
- Understanding of modern ML architectures and inference techniques
Desired Qualifications
- High agency and a track record of leading ambiguous work end to end
- Experience with AI Accelerators
- Open-source contributions / leadership
- Research in ML inference or distributed systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.