Member of Technical Staff – Engineer, Policy Evaluation & Experiment Infrastructure
$152,250–$202,750 year
On-siteSan Francisco, California, United States or Cambridge, Massachusetts, United States
Job Summary
Build and maintain pipelines that evaluate robot policies in simulation and on hardware, while developing the experiment-tracking tooling the team uses daily. Implement task-success metrics, dashboards, and automated checks to turn messy robot rollouts into trustworthy, comparable signal. Improve the speed and reliability of the experiment loop so researchers can evaluate policies faster, and track down issues across the evaluation stack and in policy rollouts—from flaky jobs to misleading metrics to surprising robot behavior. Focus on delivering end-to-end solutions with a bias toward building reliable, well-tested tooling. This high-leverage engineering role sits at the center of the research loop, offering strong mentorship alongside senior engineers as you grow ownership of the autonomy stack.
Required Qualifications
- Solid engineering fundamentals and a bias toward building reliable, well-tested tooling
- Hands-on experience working with robot policies or robotic systems—running, evaluating, or debugging them (research, projects, or industry)
- Genuine interest in ML systems and evaluation, and enthusiasm for working where research meets real robots
- The ability to take a well-scoped problem and deliver it end to end
- Eagerness to learn quickly from feedback and code review
- Must be able to go through the E-Verify process of digital verification of your employment authorization documents
Desired Qualifications
- Experience evaluating or benchmarking robot policies in simulation or on hardware
- Exposure to ML experimentation, evaluation, or data pipelines
- Familiarity with Python and with cloud or compute infrastructure
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.