Data Scientist — Agent Evaluations & Quality
On-sitePalo Alto, California, United States
Job Summary
Architect automated evaluation pipelines to measure agent quality across capabilities and product surfaces. Translate ambiguous behaviors into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks. Build representative gold datasets and regression suites covering common workflows, ambiguous requests, long-tail behavior, edge cases, and adversarial scenarios. Define and track metrics such as task success, partial completion, tool-selection accuracy, tool-use correctness, instruction adherence, factual consistency, user corrections, latency, cost, and reliability. Design deterministic graders, model-based graders, and human-review processes; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement. Analyze traces, tool calls, model outputs, user context, and production outcomes to identify root causes and build a useful failure taxonomy. Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.
Required Qualifications
- 5+ years of experience in data science, machine learning, or analytics roles building or delivering evaluation systems, metrics frameworks, or quality measurement solutions for production systems
- Demonstrated experience designing and implementing evaluation frameworks, metrics, and grading systems for ML/AI systems in production
- Production-quality Python and SQL proficiency with the ability to build automated data pipelines and analysis code at scale
- Experience designing evaluation methodologies: success criteria definition, dataset construction, metric selection, and distinguishing useful benchmarks from misleading ones
- Statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems
- Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance as product behavior evolves
- Working knowledge of LLM behavior — including model-based graders, tool use, retrieval systems, multi-step execution, partial completion, and practical failure modes
- Analytical debugging ability: connecting quantitative patterns to individual system traces and identifying failure origins across model, prompt, context, tools, data, and application logic
- Experience building dashboards, reports, and communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders
- On-site in Palo Alto, CA
- Visa sponsorship is not available
Desired Qualifications
- Experience with LLM-as-a-judge systems, calibration, and measurement of grader agreement, false positives, and false negatives
- Prior work on evaluation or benchmarking platforms for AI systems
- Experience with agentic systems, multi-step task execution, or tool-use evaluation
- Experience working on customer-facing consumer software or production ML systems with real-world user outcomes
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.