Booz Allen Hamilton logo
Booz Allen HamiltonPosted 1 month ago

Site Reliability Engineer, Lead

$99,000–$225,000 year

On-siteChantilly, Virginia, United States

Full TimeSenior LevelMasters DegreeEnterprise

Job Summary

Lead Site Reliability Engineer responsibilities include ensuring reliability, performance, scalability, and security of critical production systems across cloud and air-gapped environments. You will lead design and implementation of observability, automation, incident response, and operational best practices while partnering with DevOps, infrastructure, and security teams to improve system resilience. Daily duties involve conducting root cause analysis, capacity planning, and driving continuous improvement initiatives to support highly available services. You will build and support the site environment, evaluate system health and stability, and develop technical tools to debug deployment problems. This role strengthens security posture to safeguard national interests within a defense technology organization focused on AI and cyber solutions.

Required Qualifications

  • 8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack
  • 8+ years of experience with Linux systems administration and networking fundamentals within AWS
  • Experience with Python scripting and automation
  • Experience with Infrastructure as Code using Terraform and Terragrunt
  • Knowledge of Kubernetes administration, troubleshooting, and operations
  • TS/SCI clearance with a polygraph
  • Bachelor's degree and 8+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, DevOps Engineering, or Platform Engineering in lieu of a degree
  • Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date
  • Experience with deploying and managing OpenTelemetry
  • Experience with AWS CloudWatch, AWS EKS, and related AWS services
  • Experience managing Kubernetes environments through Rancher
  • Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management
  • Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence
  • Knowledge of distributed systems, microservices architectures, and containerized workloads
  • Knowledge of NIST 800-53 and NIST-190
  • Master's degree in a relevant field
  • Security+ CE, SSCP, CCNA-Security, or GSEC Certification
  • Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information
  • TS/SCI clearance with polygraph is required

Desired Qualifications

  • Experience with deploying and managing OpenTelemetry
  • Experience with AWS CloudWatch, AWS EKS, and related AWS services
  • Experience managing Kubernetes environments through Rancher
  • Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management
  • Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence
  • Knowledge of distributed systems, microservices architectures, and containerized workloads
  • Knowledge of NIST 800-53 and NIST-190
  • Master's degree in a relevant field
  • Security+ CE, SSCP, CCNA-Security, or GSEC Certification

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce