Umbra logo
UmbraPosted 3 weeks ago

Senior Site Reliability Engineer

$150,000–$180,000 year

On-siteArlington, Virginia, United States or Reston, Virginia, United States

Full TimeSenior LevelBachelors DegreeSmall

Job Summary

Ensure the reliability and scalability of mission-critical systems through proactive monitoring, effective incident response, and on-call rotation support. Develop and promote new technologies by conducting research, creating proofs of concept, and implementing solutions that enhance platform performance and resilience. Lead by example in fostering a culture of excellence while collaborating with cross-functional teams to align on technical strategy and drive meaningful organizational improvements. Continuously evaluate and improve team processes and workflows to increase efficiency and reduce complexity.

Required Qualifications

  • Bachelor's degree in Computer Science or a related technical field
  • 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems
  • Extensive experience with AWS services (EC2, S3, Lambda, VPC Networking) and deep knowledge of cloud infrastructure, networking, and security best practices
  • Proficiency running, optimizing, and scaling Kubernetes clusters in production environments
  • Experience using and writing Terraform to architect and manage production infrastructure
  • Ability to create and utilize Infrastructure-as-code (IaC), GitOps practices, and automation tools to increase reliability and reduce manual tasks
  • Proven success in leading teams or projects using Agile/Scrum methodologies
  • Expertise in infrastructure and software architecture, capable of designing and implementing large-scale, reliable systems with minimal guidance
  • Experience developing and managing comprehensive infrastructure monitoring and alerting strategies
  • U.S. citizen, U.S. national, U.S. lawful permanent resident, refugee or asylee as defined by 8 U.S.C. § 1324b(a)(3), or must otherwise be eligible to obtain the required authorizations from the U.S. Department of State and/or U.S. Department of Commerce

Desired Qualifications

  • 10+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems
  • Advanced understanding of cloud and application security, identity management, and compliance
  • Expertise in service mesh and service registration technologies, focusing on performance and reliability
  • Experience in the aerospace industry

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce