Together AI logo
Together AIPosted 1 month ago

Research Engineer, Post-Training Inference

$200,000–$290,000 year

On-siteSan Francisco, California, United States

Part TimeSmallArtificial Intelligence

Job Summary

Design and build systems for customizing open-source models, integrating the Model Shaping and Inference platforms to ensure a seamless path from post-training to production serving. Add features to inference engines for large-scale post-training experiments, including optimizations for RL workloads, while collaborating with product, research, and engineering teams to maintain API reliability and performance. Participate in an on-call rotation to ensure 24/7 platform availability. This role supports the foundational layer of the open-source AI ecosystem, enabling developers to efficiently create high-quality models tailored to specific applications.

Required Qualifications

  • 2+ years of experience building and deploying machine learning-based services in a production environment
  • hands-on experience with modern inference engines, such as SGLang, vLLM, and TensorRT-LLM
  • familiarity with the latest methods for fine-tuning LLMs and other AI models
  • strong software engineering background in Python or Go
  • Stay up to date with the latest advances and trends in the machine learning community

Desired Qualifications

  • Serving low-precision (FP4/FP8) models
  • multiple LoRA adapters within one model instance (Multi-LoRA)
  • models distributed across several GPU nodes
  • Optimizing the performance of RL training workloads
  • Developing CUDA/Triton/CuTE DSL kernels for inference
  • Developing large-scale and high-load production systems
  • Maintaining or contributing to open-source ML projects
  • Managing machine learning workloads on Kubernetes clusters

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce