Member of Technical Staff – Staff Engineer, Multimodal Pre-Training
$318,750–$425,000 year
On-siteSan Francisco, California, United States or Cambridge, Massachusetts, United States
Job Summary
Drive the pre-training agenda for omni-models spanning language, vision-language, video, and action in close partnership with the AI team. Lead modeling work across architecture, tokenization, and scaling-law analysis to determine compute allocation while deciding how to weight and balance multiple losses. Turn petabyte-scale text, image, and video data into training-ready mixtures with the data and infrastructure teams. Raise the bar on training reliability and evaluation, support and mentor teammates, and bring frontier ideas into the stack quickly.
Required Qualifications
- Deep, hands-on experience pre-training large models—LLM, VLM, multimodal, or VLA—at significant scale, with concrete results
- Experience training across multiple objectives or modalities—e.g. combining language, vision-language, and generative video/image losses—and balancing them in one model
- Hands-on experience with generative modeling of images and/or video (diffusion, autoregressive, or masked objectives) and/or world-model approaches to physical dynamics
- Fluency with scaling laws, distributed training, and the practical failure modes of massive runs, and the ability to design around them
- A track record of setting direction others build on, moving fluidly between research and production-grade code, and prioritizing ruthlessly under uncertainty
- Must be able to lift 50 lbs
Desired Qualifications
- Experience pre-training LLMs or VLMs at scale
- Experience training video-generation, image-generation, or world-model systems at scale
- Experience with action/robotics modalities or VLA models for embodied agents
- Contributions to widely used models, influential papers, or open-source training stacks
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.