AI Evaluation Engineer
$80,000–$180,000 year
HybridToronto, Ontario, Canada
Job Summary
Design and execute structured evaluation scenarios to validate AI agent accuracy, reliability, safety, and compliance against business rules and regulatory criteria. Conduct regression evaluations across releases to monitor behavioral drift, performance degradation, and newly introduced failure modes while documenting and tracking quality issues. Partner with Agent Engineers, Architects, and Delivery teams to support production-readiness decisions, optimize prompts and workflows, and drive continuous improvements in evaluation tooling and automation. Analyze evaluation results to identify root causes and recurring trends, ensuring AI solutions deliver consistent, trusted outcomes in regulated banking environments.
Required Qualifications
- Typically 3–10+ years of relevant experience in software quality engineering, AI evaluation, AI quality engineering, machine learning evaluation, software testing, or related disciplines
- Degree in Computer Science, Software Engineering, Data Science, Artificial Intelligence, or related discipline, or equivalent practical experience
- Strong understanding of AI evaluation, large language model behaviour, reasoning quality, hallucination detection, safety, instruction adherence, factual accuracy, and business correctness
- Experience with structured software testing, regression evaluation, production-readiness assessment, and quality engineering
- Working knowledge of SDLC, CI/CD, automated evaluation, AI observability, and engineering delivery practices
- Familiarity with benchmark management, evaluation tooling, quality automation, and AI engineering workflows
- Understanding of privacy, security, governance, and regulatory considerations relevant to enterprise AI
- Proficiency with Python, SQL, or similar tools supporting evaluation and analysis
Desired Qualifications
- Experience leading evaluation activities, mentoring technical professionals, or coordinating quality initiatives is advantageous for more senior levels
- Experience evaluating LLMs, RAG systems, AI agents, or agentic AI platforms
- Experience with AI evaluation platforms such as LangSmith, OpenAI Evals, or comparable tools
- Experience integrating automated evaluation into CI/CD or MLOps workflows
- Experience with model observability, behavioural-drift detection, or AI production monitoring
- Banking, financial services, or other regulated industry experience
- Experience leading technical teams, quality initiatives, or engineering improvement programs
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.