Zeta Global logo
Zeta GlobalPosted 1 month ago

Lead Site Reliability Engineer

On-siteBengaluru, Karnataka, India

Full TimeSenior LevelLargeMarketing Technology

Job Summary

Implement and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets to drive reliability efforts. Develop resilient systems ensuring 99.9%+ uptime, automate incident detection using runbooks, and lead blameless postmortems for root cause analysis. Design full observability with Open Telemetry, capacity planning, and performance testing to support scaling. Write software for efficiency needs, champion Infrastructure as Code with Terraform or Pulumi, and participate in chaos engineering initiatives. Collaborate with development teams on AWS, Kubernetes, and EKS infrastructure. Join the on-call rotation and apply advanced alerting and anomaly detection to metrics.

Required Qualifications

  • 3-5 years of experience as an SRE, working in cloud-based environments and on-prem environments
  • Deep understanding of Linux systems, networking, and systems administration
  • Experience with cloud platforms like AWS, with a strong understanding of Kubernetes and container orchestration tools
  • Hands-on experience with observability tools such as Honeycomb, Grafana, Prometheus, Thanos, ELK (Elastic Stack), or Loki
  • Strong skills in at least one programming language (Python, Go) to write production level code
  • Strong skills in shell scripting using bash or similar
  • Experience with OpenTelemetry or other distributed tracing systems, including tracing, metrics, and logs integration
  • Experience with Chaos Engineering methodologies and tools (Chaos Mesh, chaos monkey, AWS Fault Injection Simulator, etc)
  • Reliability-focused mindset with the ability to balance fast product iterations and system stability
  • Solid understanding of SLOs, SLIs, and error budgets
  • Hands-on knowledge of CI/CD pipelines and infrastructure automation
  • Proven expertise in incident management, postmortems, and root cause analysis
  • Knowledge of modern deployment strategies (e.g., blue-green deployments, canary releases) and resiliency patterns (circuit breakers, retry mechanisms, etc)

Desired Qualifications

  • Experience with distributed systems
  • Experience with statistical analysis applied to metrics
  • Familiarity with high-performance, low-latency systems
  • Strong problem solving skills
  • Experience as on-call engineer
  • Hands-on experience running Chaos Engineering drills and initiatives

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce