Chercheur(euse) Scientifique Appliqué(e) Senior
RemoteCanada
Job Summary
Design and build the agent execution harness that routes inputs, manages context, and invokes tools across multi-step agentic workflows. Own the runtime's fault tolerance, latency, and throughput while instrumenting the system with tracing and cost attribution to catch failures before customers. Build prompt management systems for versioning and evaluation, and design frameworks that measure agent quality and drive data-informed decisions. Integrate frontier LLMs by managing routing, fallback strategies, and cost-latency tradeoffs in production. Define system boundaries for agent logic and establish design standards through architecture decisions and code reviews.
Required Qualifications
- 4+ years building production software systems with a strong track record on reliability, performance, and scalability
- Hands-on experience shipping generative AI products — not just integrating LLM APIs or building prototypes, but owning AI-powered features that production users depend on
- Solid depth in how large language models work: failure modes, context constraints, and how prompt design shapes model behavior at scale
- Practical prompt engineering experience: systematically designing, versioning, and evaluating prompts across model updates or A/B evaluation cycles
- A real track record in eval engineering — not just familiarity, but a portfolio of evaluation suites designed, shipped, and used to drive quality decisions in production AI systems
- Cost and efficiency awareness at the system level: experience reasoning about model routing, inference cost, and latency tradeoffs in production
- Solid fundamentals in software engineering: distributed systems, API design, and test discipline
- Comfort in fast-paced, ambiguous, startup-like product environments
Desired Qualifications
- Experience with multi-agent coordination models (A2A, MCP)
- Familiarity with agent frameworks (LangChain, LlamaIndex, or similar)
- Prior experience deploying AI systems in enterprise software
- Experience with AI observability tools (tracing, cost tracking, LLM-specific monitoring)
- Familiarity with cloud-native infrastructure, service observability, logging, monitoring, SRE, and production troubleshooting
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.