Garner Health logo
Garner HealthPosted 1 week ago

Staff Site Reliability Engineer

$241,000–$270,000 year

Remote

Full TimeSenior LevelMediumHealthcare

Job Summary

Architect and own the end-to-end reliability, performance, and resilience of Garner's cloud environments on AWS and Kubernetes, including AI/ML workloads. Design the SLO framework, lead incident response programs with deep-dive root cause analysis, and build the observability platform to detect issues before users feel them. Transform ambiguous scaling requirements into automated, infrastructure-as-code deliverables using Terraform and Python, while proactively optimizing cloud cost-efficiency and performance. Serve in the on-call rotation to drive corrective actions and mentor engineers across the organization to raise operational rigor. Uphold security and HIPAA compliance standards for all infrastructure changes. This Staff SRE role sits on the Platform Engineering team, leveraging AI tools to convert manual operations into hands-free, monitored processes as the senior reliability voice in the organization.

Required Qualifications

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred), with a track record of architecting reliability for systems at scale
  • Experience designing an organization's reliability practice (SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews) and the judgment to know when to build vs. buy
  • Strong Python or Go skills applied to infrastructure automation (Kubernetes API experience a plus)
  • Track record driving cloud cost-efficiency and performance optimization across compute, storage, and networking
  • Mentorship experience and the ability to set technical direction as the senior reliability voice
  • Excellent communication skills—able to make complex reliability concepts land with both technical and non-technical stakeholders
  • Fluency with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
  • Experience supporting AI/ML or data-intensive workloads in production is a plus
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
  • Comfortable with remote work and occasional travel to HQ

Desired Qualifications

  • Experience with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
  • Experience supporting AI/ML or data-intensive workloads in production is a plus
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
  • Kubernetes API experience a plus
  • Experience with AWS preferred

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce