Research Engineer, Post-Training Inference
$200,000–$290,000 year
On-siteSan Francisco, California, United States
Job Summary
Design and build systems for customizing open-source models, integrating the Model Shaping and Inference platforms to ensure a seamless path from post-training to production serving. Add features to inference engines for large-scale post-training experiments, including optimizations for RL workloads, while collaborating with product, research, and engineering teams to maintain API reliability and performance. Participate in an on-call rotation to ensure 24/7 platform availability. This role supports the foundational layer of the open-source AI ecosystem, enabling developers to efficiently create high-quality models tailored to specific applications.
Required Qualifications
- 2+ years of experience building and deploying machine learning-based services in a production environment
- hands-on experience with modern inference engines, such as SGLang, vLLM, and TensorRT-LLM
- familiarity with the latest methods for fine-tuning LLMs and other AI models
- strong software engineering background in Python or Go
- Stay up to date with the latest advances and trends in the machine learning community
Desired Qualifications
- Serving low-precision (FP4/FP8) models
- multiple LoRA adapters within one model instance (Multi-LoRA)
- models distributed across several GPU nodes
- Optimizing the performance of RL training workloads
- Developing CUDA/Triton/CuTE DSL kernels for inference
- Developing large-scale and high-load production systems
- Maintaining or contributing to open-source ML projects
- Managing machine learning workloads on Kubernetes clusters
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.