Member of Technical Staff — Inference Infrastructure
On-siteSan Francisco, California, United States
Job Summary
Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations. Design and implement techniques that improve latency, throughput, and efficiency for real-time inference, optimizing the stack to fully utilize hardware FLOPs, bandwidth, and memory. Extend orchestration frameworks like Kubernetes, Ray, and Slurm for distributed inference and large-batch evaluation sweeps. Establish standards for reliability, observability, and reproducibility across the inference stack to ensure every evaluation is trustworthy and repeatable. Collaborate with researchers to enable high-performance inference for novel architectures as they emerge.
Required Qualifications
- Experience building or optimizing inference and serving systems for throughput and latency
- Understanding of distributed compute, GPU parallelism, and hardware-aware optimization
- Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures
- Strong engineering skills: performant, maintainable code and the ability to debug complex codebases
Desired Qualifications
- contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.