JPMorgan Chase & Co logo
JPMorgan Chase & CoPosted 3 weeks ago

Senior Lead Software Engineer - Agentic AI SRE Platform

On-siteSeattle, Washington, United States

Full TimeSenior LevelEnterpriseFinancial Services

Job Summary

Design and develop enterprise AI systems that enhance engineering productivity, system reliability, and operational excellence by architecting autonomous agents that reason, analyze, and act across telemetry pipelines and cloud infrastructure. Govern agentic system components including orchestration, memory management, and human-in-the-loop patterns while owning non-functional requirements for observability, security, and guardrails. Provision scalable AI infrastructure, manage deployment pipelines, and optimize costs through creative software solutions and technical troubleshooting. Lead evaluation sessions with external vendors to drive architectural designs and technical applicability within existing systems.

Required Qualifications

  • Formal training or certification on software engineering concepts
  • 5+ years applied experience
  • Advanced proficiency in programming with Python
  • Proficiency in all aspects of the Software Development Life Cycle
  • Proficiency in automation and continuous delivery methods
  • Hands-on experience delivering agentic AI to production
  • Experience with AI agent frameworks (e.g., LangGraph, CrewAI)
  • Experience with runtimes, orchestration, evaluation/benchmarking systems, and model selection policies
  • Experience developing production generative AI applications at scale
  • Experience with experimentation, model quality measurement, retrieval-augmented generation, and vector search
  • Strong cloud infrastructure experience (e.g., AWS)
  • Experience with infrastructure-as-code (e.g., Terraform)
  • Experience building distributed systems or high-scale systems where reliability matters

Desired Qualifications

  • Previous experience working within infrastructure, platform engineering, or reliability engineering domains
  • Knowledge of site reliability engineering concepts, including Service Level Objectives/Indicators, error budgets, incident management, resilience patterns, and observability
  • Previous experience as a machine learning software engineer in a dynamic technology company or startup

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce