Research Engineer, Evals - Member of Technical Staff
On-siteLondon, England, United Kingdom
Job Summary
Design benchmarks and evaluation suites for agentic behavior, including multi-turn, long-horizon, and tool-using tasks in non-stationary environments. Build methodologies defining statistical power, variance, and construct validity while red-teaming evaluations to ensure result quality. Transform raw traces into structured evidence through failure taxonomies and attribution of outcomes to specific decisions. Develop durable evaluation and observability infrastructure that the entire company relies on. Move from measurement to prediction by inferring system capabilities from partial evidence and estimating performance before execution.
Required Qualifications
- Evidence that you can run research of your own
- Deep hands-on experience with LLMs in agentic settings
- Statistical discipline
- Strong engineering
- Strong communication skills
Desired Qualifications
- Evaluations or benchmarks you built that other people went on to use
- Broad training in empirical method, potentially from a field outside machine learning
- Experience with automated grading and model-based judging
- A published track record in a relevant field - first-author work at venues such as NeurIPS, ICML, ICLR or ACL, including the datasets and benchmarks tracks
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.