Mainframe SRE
On-siteBengaluru, Karnataka, India
Job Summary
Lead global SRE, Operations, and Production Support teams to ensure high availability, operational excellence, and scalability of business-critical platforms across AWS, Azure, and GCP environments. Own Major Incident Management (P1/P2) processes, conduct root cause analysis, and drive corrective actions while reducing recurring incidents through automation and AIOps initiatives. Establish and track SLOs, SLIs, Error Budgets, and Reliability KPIs to achieve defined SLAs and customer outcomes. Manage operational budgets, forecast staffing requirements, and optimize support costs through process improvements. Act as the primary escalation point for customers and leadership, delivering executive-level updates on service health and building trusted relationships with Engineering, Product, and Business teams.
Required Qualifications
- SRE Principles and Practices
- DevOps Methodologies
- Cloud Platforms (AWS, Azure, GCP)
- Kubernetes and Containers
- Linux/Unix Administration
- Observability Tools: Datadog, Dynatrace, Splunk, Prometheus, Grafana, New Relic
- CI/CD Pipelines
- Infrastructure as Code (Terraform, Ansible)
- Automation & Scripting (Python, PowerShell, Bash)
- ITIL Framework and Service Governance
- Executive Communication and Stakeholder Management
- Key Metrics / KPIs
- Service Availability (%)
- SLA Achievement (%)
- MTTR and MTTD
- Incident Volume Reduction
- Automation Coverage
- Customer Satisfaction (CSAT)
- Operational Efficiency Improvements
- Error Budget Compliance
- Platform Reliability Score
Desired Qualifications
- Bachelor's degree in Computer Science, Engineering, or related field.
- 12+ years of IT Operations / Infrastructure experience.
- 5+ years leading SRE, DevOps, NOC, or Production Support teams.
- Experience managing global teams and enterprise customers.
- Preferred certifications: AWS/Azure, ITIL, Kubernetes, SRE Foundation.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.