Member of Technical Staff – Senior Engineer, Reinforcement Learning – Policy Post-Training
$255,000–$340,000 year
On-siteSan Francisco, California, United States or Cambridge, England, United Kingdom
Job Summary
Design and run reinforcement learning for contact-rich manipulation, including reward design, on/off-policy methods, and stable training at scale. Drive the RL side of sim-to-real by implementing domain randomization, system identification, and iterative transfer loops with the simulation team. Combine simulation, offline data, and real-world feedback into policies that improve predictably, while crafting curricula that prevent reward hacking. Partner with manipulation, controls, and simulation teams to carry algorithm design through to results on real hardware, and mentor MTS engineers on the RL stack.
Required Qualifications
- Strong hands-on RL experience with demonstrated results on hard, real-world, or large-scale problems
- Experience applying learning to manipulation or contact-rich control (grasping, dexterous, or bimanual manipulation)
- Demonstrated success transferring learned policies from simulation to physical robots
- Ability to design rigorous experiments and drive an ambiguous workstream independently to a result
- Strong fundamentals and comfort with large training/simulation codebases
Desired Qualifications
- Experience with dexterous hands, high-DOF manipulators, or bimanual systems
- Familiarity with GPU-accelerated simulation for manipulation (Isaac, MuJoCo, etc.)
- Background in offline RL, learning from demonstration, or combining RL with imitation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.