Staff Site Reliability Engineer- Developer Platform
$186,000–$232,500 year
RemoteUnited States
Job Summary
Design and implement infrastructure components for developer portals and tooling that reduce cognitive overhead. Manage AWS footprints by optimizing multi-region availability, disaster recovery, and cost efficiency. Develop custom Kubernetes operators and controllers to automate stateful services, while maintaining a library of Infrastructure as Code modules. Build automated security controls and policy-as-code to ensure platform compliance. Identify friction points across engineering teams to collaboratively build intuitive tools that improve developer experience. Support mid-level and junior engineers through code reviews, pair programming, and technical leadership. Work closely with engineering teams to gather requirements, identify dependencies, and provide input on risk mitigation.
Required Qualifications
- 5+ years of relevant experience in Platform Engineering, DevOps, or SRE roles
- Proficiency with Infrastructure as code, preferably Terraform
- Experience designing reusable modules and managing state at scale
- Deep understanding of Kubernetes internals, networking, storage and service mesh architectures
- Proven track record implementing GitOps workflows using ArgoCD or Flux in production environments
- Strong understanding of automation and scripting (e.g., Python, Bash, GoLang)
- Strong experience with one of the major Cloud Service Providers (AWS, Azure, or GCP), including IAM, compute, networking and storage layers
- Experience maintaining and optimizing CI pipelines
- Ability to clearly communicate technical concepts to peers and stakeholders
- Comfortable working in a collaborative environment where architectural decisions are discussed and debated
- Willingness to mentor peers and share technical knowledge
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.