Senior Site Reliability Engineer – Unified Observability
On-siteHyderabad, Telangana, India
Job Summary
Design and implement enterprise observability solutions across Azure, GCP, Kubernetes, and hybrid environments. Develop monitoring, logging, tracing, and telemetry standards using industry best practices to build dashboards providing real-time infrastructure, application, and customer health visibility. Define SLIs, SLOs, error budgets, and operational health metrics while improving proactive detection and incident response through automation. Integrate observability with ServiceNow and partner with Product Engineering, Infrastructure, Security, and Operations teams to improve platform reliability. Support initiatives involving AI-driven observability and event correlation.
Required Qualifications
- Site Reliability Engineering (SRE)
- Kubernetes (AKS/GKE)
- Azure and Google Cloud Platform
- Grafana, Datadog, Prometheus, OpenTelemetry (or similar)
- Monitoring, logging, distributed tracing, and telemetry
- Infrastructure as Code (Terraform)
- Python, Go, or PowerShell automation
- CI/CD and DevOps practices
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.