Senior Site Reliability Engineer (SRE)
On-sitePune, Maharashtra, India or Noida, Uttar Pradesh, India
Job Summary
Build, automate, and operate Hanwha Vision's global cloud platform by writing clean Go or Python code to replace manual tasks with stable systems. Manage large-scale Amazon EKS clusters, including networking and scaling via Karpenter, while maintaining modular Terraform and Helm templates. Lead troubleshooting for critical outages, write clear post-mortems, and define SLIs/SLOs using Datadog and PagerDuty. Construct secure, self-service tools to meet SOC2 requirements. Requires 10+ years of SRE experience, Go/Python proficiency, and expert Terraform skills. Night shift only. Located in Ahmedabad, India.
Required Qualifications
- 10+ years of professional experience in SRE, DevOps, or Systems/Infrastructure Engineering
- Strong programming skills in Go or Python
- Production experience running and scaling Kubernetes
- Expert-level knowledge of Terraform
- Experience with Datadog, Prometheus, Grafana, or PagerDuty
- Strong written English proficiency
Desired Qualifications
- AWS certified Solution Architect
- AWS certified DevOps
- Basic understanding of DynamoDB, Kafka/MSK, or caching layers
- Experience with Amazon EKS clusters
- Experience with Karpenter
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.