LogicMonitor logo
LogicMonitorPosted 1 week ago

Sr. Evaluation Engineer

$158,400–$217,800 year

HybridSan Francisco, California, United States

Full TimeSenior LevelLarge

Job Summary

Design and build production-grade evaluation pipelines, golden datasets, automated graders, and regression frameworks for AI agents, retrieval systems, and complex investigation workflows. Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness. Build offline and online evaluation systems in Python, integrate them with CI/CD, and calibrate LLM-based graders against expert human judgment. Monitor AI quality and behavioral drift in production, converting failures and customer feedback into new tests and safeguards. Establish evaluation-driven development practices and mentor other engineers.

Required Qualifications

  • 5+ years of experience in software engineering, machine learning, applied AI, or a related field
  • Strong Python engineering skills and experience building production systems
  • Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks
  • Experience using multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms
  • Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating
  • Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering
  • Experience evaluating non-deterministic, multi-step, or multi-agent AI systems
  • Ability to translate human and domain-expert judgment into test cases, evaluation rubrics, scoring functions, and automated graders
  • Experience with LLM-as-a-judge techniques, including grader design, calibration, reliability measurement, and alignment with expert human judgment
  • Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis
  • Strong analytical, systems-thinking, and communication skills
  • Residents of California
  • Candidates who currently hold valid U.S. work authorization that can be transferred to a new employer (such as certain H-1B statuses)
  • Candidates authorized to work in the United States on a full-time, permanent basis without requiring new or initial employer-sponsored work authorization

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce