Senior Site Reliability Engineer - Managed Kubernetes
$240,000–$356,000 year
HybridSan Francisco, California, United States or San Jose, California, United States
Job Summary
Operate and maintain bare-metal Kubernetes clusters scaling to thousands of nodes, handling degradation, recovery, resizing, and incident response via fleet management tools. Design, build, and maintain scalable control plane services, operators, and custom controllers while developing automation for cluster lifecycle management including provisioning, upgrades, and patching. Assist customers with Kubernetes questions regarding workload integration, storage, and authentication, and participate in a well-managed on-call rotation for critical incidents. Use Python and Golang to create tooling that automates the validation of platform quality and define SLOs and SLIs for Kubernetes services and platform reliability.
Required Qualifications
- 6+ years of experience in a SRE, operations engineer, or similar role
- Deep knowledge of running Linux clusters and systems
- Strong programming skills in Go and Python
- Experience with GitOps (e.g., ArgoCD), Helm, and Kubernetes operators
- Proven experience operating Kubernetes clusters in production environments (on-prem, EKS, GKE, or similar)
- Familiarity with observability tools like Prometheus, Grafana, FluentBit, and CI/CD pipelines
- Proven experience provisioning Kubernetes using tools such as kubeadm, Cluster API, or similar
- Presence in our San Francisco, San Jose, or Bellevue office location 4 days per week
Desired Qualifications
- Deep Kubernetes expertise: CRDs, CSI, CNI, Kubernetes Operator Coding experience
- Exposure to HPC clusters, AI/ML workloads, or large-scale GPU clusters
- Hybrid or multi-cloud Kubernetes environment experience
- Contributions to CNCF projects or Kubernetes SIGs
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.