Zafin logo
ZafinPosted 1 month ago

AI Evaluation Engineer

$80,000–$180,000 year

HybridToronto, Ontario, Canada

Full TimeMediumBanking and Financial Services

Job Summary

Design and execute structured evaluation scenarios to validate AI agent accuracy, reliability, safety, and compliance against business rules and regulatory criteria. Conduct regression evaluations across releases to monitor behavioral drift, performance degradation, and newly introduced failure modes while documenting and tracking quality issues. Partner with Agent Engineers, Architects, and Delivery teams to support production-readiness decisions, optimize prompts and workflows, and drive continuous improvements in evaluation tooling and automation. Analyze evaluation results to identify root causes and recurring trends, ensuring AI solutions deliver consistent, trusted outcomes in regulated banking environments.

Required Qualifications

  • Typically 3–10+ years of relevant experience in software quality engineering, AI evaluation, AI quality engineering, machine learning evaluation, software testing, or related disciplines
  • Degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or related discipline, or equivalent practical experience
  • Strong understanding of AI evaluation, large language model behaviour, reasoning quality, hallucination detection, safety, instruction adherence, factual accuracy, and business correctness
  • Experience with structured software testing, regression evaluation, production-readiness assessment, and quality engineering
  • Working knowledge of SDLC, CI/CD, automated evaluation, AI observability, and engineering delivery practices
  • Familiarity with benchmark management, evaluation tooling, quality automation, and AI engineering workflows
  • Understanding of privacy, security, governance, and regulatory considerations relevant to enterprise AI
  • Proficiency with Python, SQL, or similar tools supporting evaluation and analysis

Desired Qualifications

  • Experience leading evaluation activities, mentoring technical professionals, or coordinating quality initiatives is advantageous for more senior levels
  • Experience evaluating LLMs, RAG systems, AI agents, or agentic AI platforms
  • Experience with AI evaluation platforms such as LangSmith, OpenAI Evals, or comparable tools
  • Experience integrating automated evaluation into CI/CD or MLOps workflows
  • Experience with model observability, behavioural-drift detection, or AI production monitoring
  • Banking, financial services, or other regulated industry experience
  • Experience leading technical teams, quality initiatives, or engineering improvement programs

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce