Senior Site Reliability Engineer - Python | Go | Kubernetes | DevOps | Bangalore | 8-10 Yrs
On-siteBengaluru, Karnataka, India
Job Summary
Operate and improve Kubernetes-based production infrastructure and deployment systems for the Splunk Agent Resilience team, owning customer deployments across cloud and air-gapped environments including installation, upgrades, troubleshooting, and lifecycle management. Build and enhance deployment observability, monitoring, logging, and alerting while developing automation to improve operational efficiency. Participate in production incident response, root cause analysis, and reliability improvements, tuning infrastructure components such as databases and services to boost performance. Design and develop internal tooling using Python and/or Go, and manage infrastructure using Terraform or similar Infrastructure as Code tools. Collaborate with software engineers and customers to design secure, scalable, and reliable deployment architectures.
Required Qualifications
- 8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Infrastructure Engineering, or a related field
- 3+ years' operating Kubernetes in production
- experience with Helm
- Experience building and maintaining CI/CD platforms and deployment automation
- Experience working with AWS, GCP, or similar cloud platforms
Desired Qualifications
- Experience improving production reliability, scalability, and availability
- Experience with monitoring, logging, observability, and alerting platforms
- Strong scripting or programming skills in Python and/or Go
- Experience with Infrastructure as Code tools such as Terraform
- Experience with MLOps
- Solid understanding of networking fundamentals (VPCs, DNS, routing, load balancing)
- Experience operating and tuning data infrastructure for performance and reliability — OLTP/OLAP databases, message queues, and object storage systems
- Experience deploying and supporting cloud and air-gapped/on-prem environments
- Strong debugging and troubleshooting skills across distributed systems
- Ability to collaborate effectively across engineering, infrastructure, security, and customer-facing teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.