JPMorgan Chase & Co logo
JPMorgan Chase & CoPosted 1 week ago

Senior Lead Software Engineer- AI/ML Platform

On-siteWilmington, Delaware, United States

Full TimeSenior LevelEnterpriseFinancial Services

Job Summary

Build and maintain reusable AI/ML platform infrastructure and shared services to support development, deployment, and operations at scale. Architect, deploy, and operate secure cloud and container-based environments for training and inference, including GPU-intensive workloads. Design and implement platform tooling, automation, and infrastructure-as-code solutions to streamline model deployment, environment provisioning, release management, and operational support. Develop production-grade services, APIs, SDK integrations, and workflows that support model training, serving, evaluation pipelines, and AI application lifecycle management. Partner with data science, ML engineering, and application teams to translate model and compute requirements into platform standards and deployment patterns. Optimize platform reliability, scalability, latency, and cost through orchestration, scheduling, and hardware acceleration. Establish operational best practices including monitoring, logging, observability, access controls, incident response, and production troubleshooting. Drive adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes.

Required Qualifications

  • Formal training or certification on software engineering concepts
  • 5+ years applied experience
  • Experience delivering secure, production-quality code in Python or Java
  • Strong foundations in distributed systems, microservices, and platform architecture/design principles
  • Proven ability to architect and operate cloud-native infrastructure on AWS (compute, networking, storage, security) and other major clouds
  • Demonstrated expertise with infrastructure-as-code tooling, specifically Terraform, in large-scale cloud environments
  • Hands-on experience with Docker and Kubernetes, including AWS EKS operations
  • Experience building or supporting production AI/ML platforms (training, deployment, and model serving/inference), including GPU infrastructure/tooling
  • Strong DevOps/platform engineering practices: CI/CD, release automation, automated testing, and observability (monitoring/logging/tracing)
  • Experience with SQL/NoSQL databases and data integration
  • Strong Linux, scripting, and networking fundamentals
  • Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations
  • experience coaching senior engineers/leads on compliant usage patterns and controls

Desired Qualifications

  • Proficiency in Go or Python for automation, tooling development, or platform service implementation
  • Experience with MLOps frameworks and tools such as Kubeflow, MLflow, or similar AI/ML lifecycle management platforms
  • Working knowledge of ML frameworks (PyTorch, TensorFlow, Hugging Face, scikit-learn) for model integration and operationalization
  • Exposure to multi-cloud or hybrid cloud architectures and platform portability strategies

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce