Site Reliability Engineer II
$103,500–$150,000 year
HybridMcLean, Virginia, United States
Job Summary
Collaborate with software engineering teams to improve application reliability and scalability while operating and supporting production services across Kubernetes environments. Troubleshoot infrastructure and application issues across the full technology stack, build automation and tooling to reduce operational overhead, and leverage AI-assisted engineering to accelerate troubleshooting and eliminate manual work. Monitor system health using observability platforms, participate in incident response and root cause analysis, and develop infrastructure-as-code configurations to streamline deployment workflows. This hybrid role requires a rotating on-call schedule supporting global production systems, with candidates based in the Tysons vicinity prioritized for three days per week onsite.
Required Qualifications
- 2+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Cloud Operations, or related roles
- Demonstrated experience supporting production environments running on Kubernetes or other containerized platforms
- Demonstrated experience with cloud infrastructure platforms such as AWS, OCI, or GCP
- Demonstrated experience with Linux systems administration and troubleshooting
- Demonstrated experience with scripting or programming languages such as Python, Bash, or Go
- Familiarity with CI/CD pipelines and Git-based workflows
- Demonstrated understanding of networking fundamentals including DNS, load balancing, TLS/SSL, and routing concepts
- Demonstrated experience troubleshooting distributed systems and production incidents
- Ability to participate in an on-call rotation supporting production systems
- Fluency in English, both oral and written
- Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite
Desired Qualifications
- Experience with GitOps and tools such as ArgoCD
- Experience with infrastructure-as-code tools such as Terraform
- Familiarity with observability platforms such as Prometheus, Grafana, Loki, or OpenTelemetry
- Experience operating services in hybrid-cloud or multi-region environments
- Understanding of release strategies such as rolling deployments, canary releases, or blue/green deployments
- Familiarity with incident management and operational best practices
- Exposure to security and compliance concepts in production environments
- Experience using AI-assisted development, automation, or operational tooling to improve engineering productivity and service reliability
- Demonstrated passion for automation, process improvement, and operational efficiency
- Strong communication and collaboration skills
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.