Site Reliability Engineer
$151,500–$252,500 year
RemoteUnited States or Seattle, Washington, United States
Job Summary
Design infrastructure for high availability and fault tolerance on Azure Government, define SLIs, SLOs, and error budgets, and lead end-to-end incident response with blameless postmortems. Build runbooks, architecture docs, and operational guides to establish reliability practices from the ground up while closing observability gaps through instrumentation and alerting standards. Develop automation for fleet management and release validation pipelines using IaC and CI/CD tools, adapting chaos engineering and monitoring to meet strict regulatory constraints. Collaborate with product, security, and compliance teams to identify risks and drive practical remediation plans within air-gapped environments. Participate in on-call rotations and mentor engineers to spread SRE practices across the organization.
Required Qualifications
- 7+ years in Software Engineering
- 3+ years in SRE, Platform Engineering, or similar
- Experience with Government or Sovereign Cloud (e.g., Azure Government, AWS GovCloud)
- Experience in regulated compliance environments — government (FedRAMP, CMMC, IL2/IL4/IL5), financial (PCI-DSS, SOX), or healthcare (HIPAA, HITRUST)
- Strong experience building and running production services on cloud infrastructure (Azure preferred, including Azure Government)
- Able to learn large, complex platforms quickly with limited guidance
- Can investigate systems independently and produce clear docs, risk assessments, and improvement plans
- Comfortable working across teams — engineering, product, security, compliance, operations
- Programming skills in one or more: TypeScript/JS, Go, Java, C#, or similar
- Experience with monitoring and observability tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK stack)
- Experience with IaC (Terraform, Terragrunt, Pulumi) and container orchestration (Kubernetes)
- Experience with CI/CD and GitOps tooling — GitHub Actions, Azure DevOps, GitLab CI, ArgoCD, FluxCD, or Dagger
- Solid grasp of distributed systems, networking, and cloud-native architecture
- Clear written and verbal communication skills
- Must be able to lift 50 lbs
Desired Qualifications
- Experience on B2B SaaS platforms in regulated or government markets
- Background in chaos engineering, resilience testing, or performance/load testing
- Have built an SRE or reliability function from scratch before
- Experience across mixed environments — modern cloud-native and older legacy systems
- Familiar with AI-first development workflows — using LLM-powered tools for infrastructure automation, code generation, and documentation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.