Research Scientist (Remote/US/LATAM)
RemoteUnited Kingdom or Mexico
United Kingdom or MexicoRemoteContractDoctorate Or Professional DegreeSmall
ContractDoctorate Or Professional DegreeSmall
Job Summary
Design frontier-grade evaluation packages for LLMs across reasoning, coding, and multi-modal domains. Construct original benchmark designs with expert-verified ground truth, rigorous QC, and calibrated rubrics to measure model capability. Recruit and calibrate expert pools, acting as the final arbiter of correctness and frontier difficulty. Translate lab measurement requests into winning sample packages, managing end-to-end pilots. Support publishing at top venues like NeurIPS, ICLR, and ACL. Own the full research loop from framing questions to writing papers.
Required Qualifications
- Research background in ML evaluation or benchmarking
- Track record of published or open benchmarks
- Eval/measurement research experience
- Equivalent hands-on work that labs have relied on
- Deep LLM/frontier-model benchmarking expertise
- Real strength in code-model evaluation
- Real strength in agentic evaluation
- Fluency with the measurement problem
- Knowledge of construct validity
- Knowledge of psychometrics
- Knowledge of rubrics
- Knowledge of pass rates
- Knowledge of headroom
- Knowledge of contamination
- Knowledge of what makes a task genuinely discriminate a model
- Proven ability to hold a team or expert pool to a rigorous standard
- Comfort with the full research loop
- Ability to frame the question
- Ability to run the study
- Ability to write it up
- Fluent English
Desired Qualifications
- Interest in the safety side of evaluation
- Interest in capability elicitation
- Interest in robustness
- Interest in measuring the things that are hardest to measure honestly
- Spanish
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.