Office Hours logo
Office HoursPosted 1 week ago

Software Engineer- Benchmarking

$160,000–$210,000 year

RemoteSan Francisco, California, United States

Full TimeSmall

Job Summary

Prepare and maintain benchmark datasets by cleaning, converting, and validating tasks for consistency. Build and maintain evaluation pipelines that run consistently across model APIs and terminal agents. Create lightweight, containerized environments for tasking and evaluating model performance on tool use. Develop the scoreboard and leaderboard to publish results at model, benchmark, task, domain, and rubric levels. Build analysis tools to identify failure modes and track capability improvements. Collaborate with researchers to ensure evaluation data is accurate and integrated into published outputs.

Required Qualifications

  • 4+ years of professional experience building and maintaining complex systems
  • Strong Python
  • Experience preparing, cleaning, and maintaining datasets
  • Experience with Docker and building reproducible execution environments
  • Ability to work well alongside researchers and scientists
  • Ability to translate methodology into working systems

Desired Qualifications

  • Hands-on experience running AI evaluations
  • Experience with frameworks like Harbor, Terminal-Bench, or Inspect
  • Experience fine-tuning or post-training open-source LLMs
  • Experience with agentic, multi-turn, long-context, or tool-use evaluation
  • Experience validating LLM-as-judge or rubric-based grading setups
  • Background or strong interest in a scientific or technical domain
  • Experience building data-heavy dashboards, leaderboards, or visualizations
  • Open-source contributions or published work related to benchmarks and measurement

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce