Site Reliability Engineer (SRE) / Service Availability Manager
On-siteBethesda, Maryland, United States
Job Summary
Lead incident command during major outages to restore services and minimize business impact, while designing automation tools to reduce manual intervention and improve system performance. Maintain comprehensive monitoring frameworks to detect anomalies early and conduct thorough post-mortems to identify root causes and implement preventative measures. Collaborate with development and operations teams to enhance incident response processes and communicate status updates to stakeholders at all levels. Cover shifts in a 24x7x365 environment with on-call responsibilities to ensure continuous service availability.
Required Qualifications
- 5+ years of experience in an information technology environment
- 3 years of experience in information technology focused on IT Operations that include troubleshooting complex network, server, storage, and/or application issues
- 3 years minimum operations experience involving incident, problem, change, and release management that included leading calls and documenting outcomes
- Undergraduate degree or or equivalent experience/certification
- Ability to cover shifts in a 24x7x365 environment and on-call responsibilities
- Poficiency in scripting languages (Python, Shell) and familiarity with automation tools (such as Ansible, Jenkins)
- Ability to use artificial intelligence to complete IT Operations tasks and activities
- Knowledge of cloud platforms (AWS, Azure, GCP), infrastructure as code, and containerization technologies
- Strong leadership qualities, including decisiveness, and the ability to motivate teams, along with the ability to manage stressful situations calmly and effectively
- Strong experience in incident command or incident management in a technology environment
- Excellent problem-solving, organizational, and analytical skills
Desired Qualifications
- ITIL Foundations v3+ Certification
- Demonstrated experience with ITSM suites, e.g., ServiceNow
- Demonstrated experience with various monitoring, performance, or capacity tools
- Experience with continuous integration/continuous deployment (CI/CD) pipelines and DevOps practices
- Ability to create constructive relationships, influence, and communicate with varying levels of associates and management
- Ability to solve complex, cross-functional issues
- Strong knowledge of Server, Storage, Network, Middleware, Application and Cloud technologies
- A high degree of curiosity and a drive to seek more efficient ways of delivering service
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.