Member of Technical Staff – Staff Engineer, Multimodal Post-Training
$318,750–$425,000 year
On-siteSan Francisco, California, United States or Cambridge, Massachusetts, United States
Job Summary
Shape post-training strategy for SFT, preference tuning, distillation, and RL to convert strong base models into deployable robot policies. Drive grounded reasoning, tool use, and multimodal understanding for models acting in the real world, while aligning behavior with product and safety requirements. Strengthen evaluation methodology to separate real gains from noise across offline metrics and on-robot performance. Set the technical bar for post-training, mentor the team, and partner with data teams on data that improves behavior. Requires deep hands-on experience with frontier-scale multimodal LLMs or VLAs and a record of owning end-to-end model-improvement programs.
Required Qualifications
- Deep hands-on experience post-training large multimodal models—SFT, RLHF/DPO-style preference methods, distillation—with results you have shipped or published.
- Strong command of RL for large models (reward modeling, policy optimization) and its practical failure modes.
- Experience working with multimodal LLMs or VLA models at frontier scale.
- A record of owning end-to-end model-improvement programs and setting technical direction others follow.
- Excellent judgment about what to measure, and the instinct to debug across the stack.
Desired Qualifications
- Experience with embodied AI, robotics, or vision-language-action models (or strong desire to move from LLMs into the physical world).
- Familiarity with evaluating open-ended, real-world task performance.
- Contributions to influential models, papers, or open-source post-training stacks.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.