Manager - Production Operations & Site Reliability Engineering
$140,250–$181,500 year
On-siteLake Forest, California, United States
Job Summary
Lead steady-state operations for cloud-based healthcare platforms, ensuring high availability, reliability, and performance through the Steady State Operations Framework. Drive release governance by managing production readiness reviews, validating deployment readiness across Validation, Staging, and Production environments, and enforcing standardized change success protocols. Champion SRE best practices to reduce operational toil, lead incident investigations, and improve service reliability by lowering MTTD and MTTR while increasing deployment success rates. Provide technical leadership across AWS platforms including EKS, EC2, RDS, and Kubernetes, while defining enterprise observability strategies using Datadog and distributed tracing. Ensure full compliance with HIPAA, GDPR, and FDA standards through rigorous security architecture and operational controls. Serve as the senior escalation point for major production events and mentor engineering teams on operational excellence.
Required Qualifications
- Bachelor's Degree or Equivalent years of directly related experience (or high school +13 yrs; Assoc.+9 yrs; M.S.+2 yrs; PhD+0 yrs)
- The ability to fluently read, write, understand and communicate in English
- 5 Years of Relevant Experience
Desired Qualifications
- Bachelor's Degree in degree in Computer Science, Engineering, or related field (Master's preferred)
- Experience in Production Operations, Site Reliability Engineering, Platform Engineering, or Cloud Operations
- Experience supporting regulated healthcare or medical device platforms
- AWS Solutions Architect or DevOps Professional certification
- Kubernetes (CKA/CKS) certification
- Experience with healthcare interoperability standards (HL7, FHIR, DICOM)
- Experience implementing Site Reliability Engineering practices within large-scale cloud environments
- AWS (EKS, EC2, RDS, S3, Route53, IAM, Load Balancers, CloudWatch)
- Kubernetes, Istio, Docker
- Datadog, APM, logging, distributed tracing, synthetic monitoring
- CI/CD, Infrastructure as Code, automation
- HIPAA, GDPR, FDA compliance
- Strong understanding of cloud networking, security, and high-availability architectures
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.