Senior Manager, Site Reliability Engineering
On-siteReston, Virginia, United States
Job Summary
Architect infrastructure and service reliability while supervising team members to forecast demand and identify resource gaps. Monitor data collection, triage incidents, and perform root cause analyses to ensure systems meet service level objectives and SLAs. Implement standards for automation, testing, and decommissioning to enhance operational efficiency and remove unused objects. Collaborate with software development teams to build scalable infrastructures and guide technical communication regarding service attributes. Evaluate cutting-edge tools to optimize performance, manage prioritization of improvement initiatives, and conduct post-mortems to prevent incident reoccurrence.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.