Clera logo
CleraPosted 1 week ago
EXPIRED

Data Scientist — Agent Evaluations & Quality

On-sitePalo Alto, California, United States

Full TimeDoctorate Or Professional DegreeStartup

Job Summary

Architect automated evaluation pipelines to measure agent quality across capabilities and product surfaces. Translate ambiguous behaviors into explicit success criteria, including pass, partial-pass, and failure definitions for complex multi-step tasks. Build representative gold datasets and regression suites covering common workflows, ambiguous requests, long-tail behavior, edge cases, and adversarial scenarios. Define and track metrics such as task success, partial completion, tool-selection accuracy, tool-use correctness, instruction adherence, factual consistency, user corrections, latency, cost, and reliability. Design deterministic graders, model-based graders, and human-review processes; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement. Analyze traces, tool calls, model outputs, user context, and production outcomes to identify root causes and build a useful failure taxonomy. Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

Required Qualifications

  • 5+ years of experience in data science, machine learning, or analytics roles building or delivering evaluation systems, metrics frameworks, or quality measurement solutions for production systems
  • Demonstrated experience designing and implementing evaluation frameworks, metrics, and grading systems for ML/AI systems in production
  • Production-quality Python and SQL proficiency with the ability to build automated data pipelines and analysis code at scale
  • Experience designing evaluation methodologies: success criteria definition, dataset construction, metric selection, and distinguishing useful benchmarks from misleading ones
  • Statistical and experimental design knowledge: sampling, variance, uncertainty quantification, bias detection, confounding variables, and significance testing for non-deterministic systems
  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance as product behavior evolves
  • Working knowledge of LLM behavior — including model-based graders, tool use, retrieval systems, multi-step execution, partial completion, and practical failure modes
  • Analytical debugging ability: connecting quantitative patterns to individual system traces and identifying failure origins across model, prompt, context, tools, data, and application logic
  • Experience building dashboards, reports, and communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders
  • On-site in Palo Alto, CA
  • Visa sponsorship is not available

Desired Qualifications

  • Experience with LLM-as-a-judge systems, calibration, and measurement of grader agreement, false positives, and false negatives
  • Prior work on evaluation or benchmarking platforms for AI systems
  • Experience with agentic systems, multi-step task execution, or tool-use evaluation
  • Experience working on customer-facing consumer software or production ML systems with real-world user outcomes

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Find similar roles