Software Engineer- Benchmarking
$160,000–$210,000 year
RemoteSan Francisco, California, United States
Job Summary
Prepare and maintain benchmark datasets by cleaning, converting, and validating tasks for consistency. Build and maintain evaluation pipelines that run consistently across model APIs and terminal agents. Create lightweight, containerized environments for tasking and evaluating model performance on tool use. Develop the scoreboard and leaderboard to publish results at model, benchmark, task, domain, and rubric levels. Build analysis tools to identify failure modes and track capability improvements. Collaborate with researchers to ensure evaluation data is accurate and integrated into published outputs.
Required Qualifications
- 4+ years of professional experience building and maintaining complex systems
- Strong Python
- Experience preparing, cleaning, and maintaining datasets
- Experience with Docker and building reproducible execution environments
- Ability to work well alongside researchers and scientists
- Ability to translate methodology into working systems
Desired Qualifications
- Hands-on experience running AI evaluations
- Experience with frameworks like Harbor, Terminal-Bench, or Inspect
- Experience fine-tuning or post-training open-source LLMs
- Experience with agentic, multi-turn, long-context, or tool-use evaluation
- Experience validating LLM-as-judge or rubric-based grading setups
- Background or strong interest in a scientific or technical domain
- Experience building data-heavy dashboards, leaderboards, or visualizations
- Open-source contributions or published work related to benchmarks and measurement
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.