Senior Site Reliability Engineer
$150,000–$180,000 year
On-siteArlington, Virginia, United States or Reston, Virginia, United States
Job Summary
Ensure the reliability and scalability of mission-critical systems through proactive monitoring, effective incident response, and on-call rotation support. Develop and promote new technologies by conducting research, creating proofs of concept, and implementing solutions that enhance platform performance and resilience. Lead by example in fostering a culture of excellence while collaborating with cross-functional teams to align on technical strategy and drive meaningful organizational improvements. Continuously evaluate and improve team processes and workflows to increase efficiency and reduce complexity.
Required Qualifications
- Bachelor's degree in Computer Science or a related technical field
- 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems
- Extensive experience with AWS services (EC2, S3, Lambda, VPC Networking) and deep knowledge of cloud infrastructure, networking, and security best practices
- Proficiency running, optimizing, and scaling Kubernetes clusters in production environments
- Experience using and writing Terraform to architect and manage production infrastructure
- Ability to create and utilize Infrastructure-as-code (IaC), GitOps practices, and automation tools to increase reliability and reduce manual tasks
- Proven success in leading teams or projects using Agile/Scrum methodologies
- Expertise in infrastructure and software architecture, capable of designing and implementing large-scale, reliable systems with minimal guidance
- Experience developing and managing comprehensive infrastructure monitoring and alerting strategies
- U.S. citizen, U.S. national, U.S. lawful permanent resident, refugee or asylee as defined by 8 U.S.C. § 1324b(a)(3), or must otherwise be eligible to obtain the required authorizations from the U.S. Department of State and/or U.S. Department of Commerce
Desired Qualifications
- 10+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems
- Advanced understanding of cloud and application security, identity management, and compliance
- Expertise in service mesh and service registration technologies, focusing on performance and reliability
- Experience in the aerospace industry
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.