Site Reliability Engineer
On-siteBengaluru, Karnataka, India
Job Summary
Own the health of Sigma, MiQ's enterprise platform by building world-class observability across services and infrastructure. Design monitoring, alerting, and synthetic checks to proactively surface application errors and performance regressions, then partner with product teams to drive resolutions. Work hands-on with Grafana, AWS, and Kubernetes to define SLIs/SLOs, reduce alert noise, and strengthen incident response practices. Improve release safety by optimizing pipelines, enabling progressive delivery, and implementing health gates and automated rollbacks. Apply a performance engineering mindset through load testing and capacity analysis to automate away toil. Accelerate SRE maturity by leveraging AIOps capabilities to improve signal quality and reduce manual operational effort. Over time, grow into owning the observability charter, setting standards, and evangelizing best practices across engineering teams.
Required Qualifications
- 4 - 8 years of experience in Site Reliability Engineering, DevOps, or platform/production engineering roles supporting customer-facing systems
- Hands-on experience with observability tooling — Grafana, Prometheus, Datadog, and log/trace aggregation (e.g., Loki, Tempo, OpenTelemetry, or equivalent) — covering metrics, logs, traces, and events
- Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
- Deep knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
- Solid experience operating workloads on Kubernetes (ideally EKS) and AWS — comfortable debugging issues across the application, container, and infrastructure layers
- Strong scripting and automation skills in Python, Bash, or Go, with exposure to infrastructure-as-code (e.g., Terraform) and CI/CD pipelines
- A performance engineering orientation: load/stress testing (e.g., k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
- Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
- Experience contributing to incident management — triage, escalation, communication, and post-mortems — and helping set up or improve on-call processes and rotations
- A proactive, ownership-driven mindset: you identify problems from telemetry before they're reported, and you follow through with the teams who need to fix them
- Clear written and verbal communication — you can turn noisy signals into crisp findings, runbooks, and recommendations for engineering teams
- Strong communication and collaboration abilities to work effectively alongside peer Platform teams, such as Cloud, DevOps, and Security
Desired Qualifications
- Experience with Kubernetes (ideally EKS)
- Experience with AWS
- A performance engineering orientation: load/stress testing (e.g., k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
- Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.