Staff Site Reliability Engineer, GovCloud
$158,500–$230,000 year
HybridMcLean, Virginia, United States
Job Summary
Design and operate highly available, secure cloud infrastructure on AWS GovCloud, including networking, identity management, Kubernetes clusters, and PostgreSQL. Manage end-to-end VPC architecture, load balancing, and DNS while tuning production databases for high availability and performance. Develop Infrastructure-as-Code workflows using Terraform and GitOps to deploy changes safely, and lead incident response under pressure. Partner with engineering and security teams to improve observability, reduce alert noise, and drive reliability improvements across the platform. Mentor teammates and document operational procedures to raise engineering standards as the GovCloud team scales.
Required Qualifications
- Must reside in the United States and be legally authorized to work in the US without sponsorship
- Bachelor's degree or equivalent experience in Computer Science or a related field
- 8+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles — or 5+ years with demonstrated Staff-level scope (technical leadership, cross-team delivery, incident ownership, and platform/IaC ownership)
- Production experience with: Kubernetes
- Production experience with: AWS core services (IAM, compute, object storage, encryption/key management) and AWS cloud networking (VPC design, routing, security groups, load balancing, VPC endpoints/PrivateLink, Transit Gateway, DNS)
- Production experience with: Terraform or comparable infrastructure-as-code tools
- Production experience with: Git and CI/CD pipelines
- Production experience with: Linux and foundational systems concepts (networking, DNS, TLS/certificates)
- Production experience with: PostgreSQL (or comparable relational databases) in production — replication, backups, and performance tuning
- Programming and Automation: Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services
- Incident & Change Management: Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes
- Experience participating in a production on-call rotation
- Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems
- Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite
Desired Qualifications
- Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments
- Familiarity with security and compliance practices such as FIPS, vulnerability management, and controlled production change processes
- Experience with observability and logging platforms in enterprise production environments
- Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with data/platform technologies such as Redis and Kafka
- Experience supporting federal agencies or public-sector customers
- Exposure to government networking and security requirements
- Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise
- Excellent collaboration skills and a strong willingness to learn
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.