Lead Software Engineer - AI Platform Reliability
On-siteSeattle, Washington, United States
Job Summary
Design and implement solutions to enhance the reliability and scalability of AI/ML platforms and applications. Develop secure, stable production code, participate in code reviews, and debug defects across AI Foundation Services components. Build reusable platform services, APIs, SDKs, and libraries that standardize model hosting and inference consumption. Partner with Lines of Business teams to implement AI capabilities from technical design through early operational support. Own non-functional requirements, establish standards for observability and security, and enforce reference architectures for operational readiness. Participate in on-call rotations to resolve complex production issues and drive durable remediation. Mentor engineers and raise standards for engineering quality and operational rigor.
Required Qualifications
- Formal training or certification on software engineering concepts
- 5+ years applied experience
- Strong hands-on coding experience in Python
- Experience delivering production-grade services
- Demonstrated experience leading effective use of approved AI-assisted software development tools
- Ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows
- Experience coaching engineers on safe, compliant adoption within delivery practices
- Hands-on practical experience with system design
- Experience with automated testing
- Experience with debugging
- Experience with operational stability for production software
- Experience implementing observability
- Experience implementing logging
- Experience implementing metrics
- Experience implementing alerts
- Experience implementing Service Level Objectives
- Experience implementing incident response practices
- Experience implementing root-cause analysis for services in production
- Working knowledge of software application development and technical processes
- Depth in cloud platforms, artificial intelligence, machine learning platforms, distributed systems, or infrastructure engineering
- Ability to break down technical requirements into executable engineering tasks
- Ability to manage dependencies
- Ability to deliver against milestones in partnership with product and application teams
- Strong written and verbal communication skills
- Ability to explain technical decisions, trade-offs, issues, and risks to engineering teams and stakeholders
Desired Qualifications
- Experience supporting AI/ML or generative AI platform capabilities
- Experience with model hosting
- Experience with inference services
- Experience with model gateways
- Experience with managed AI services
- Experience with developer-facing AI/ML infrastructure
- Proven skills in managing AI infrastructure on cloud platforms
- Experience with deployment of machine learning workloads
- Experience with scaling of machine learning workloads
- Experience with monitoring of machine learning workloads
- Experience with optimizing machine learning workloads
- Experience building reusable golden path assets
- Experience with templates
- Experience with reference implementations
- Experience with SDKs
- Experience with automated tests
- Experience with onboarding guides
- Experience with deployment patterns
- Experience developing generative AI applications
- Experience developing AI agents
- Experience implementing AI-assisted operations
- Experience implementing appropriate guardrails
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.