Cloud Site Reliability Senior Engineer
On-siteBengaluru, Karnataka, India
Job Summary
Lead reliability improvements for production services by driving troubleshooting, incident response, and RCA across cloud services. Partner with Engineering and Operations to enhance release readiness, deployment practices, and operational standards while strengthening observability and alerting. Support cloud platform improvements, automation, and resilience initiatives using Terraform-managed infrastructure on Azure and AWS to reduce toil. Strengthen operational insights and share knowledge to reduce silos across SRE and Cloud Operations teams. Contribute to secure, compliant, and cost-conscious operations for containerized services and AKS environments.
Required Qualifications
- Bachelor's degree in computer science engineering, Information Technology, or equivalent degree
- 5+ years of progressive experience in Site Reliability Engineering, Cloud Operations, DevOps, or Platform Engineering
- Strong Linux/Unix command-line administration, troubleshooting, package management, and systems operations skills
- Experience managing production cloud infrastructure across Azure and/or AWS environments
- Hands-on experience operating container orchestration platforms, including AKS migrations, upgrades, and production operations
- Hands-on Infrastructure as Code experience, preferably with Terraform for provisioning and lifecycle management
- Automation, scripting, and AI-assisted engineering experience using Python, Bash, Go, YAML, generative AI tools, or similar technologies
- Experience with CI/CD pipelines using Azure DevOps, ArgoCD, Jenkins, GitHub Actions, or similar platforms
- Experience with Docker, container registries such as Azure Container Registry, and container lifecycle management
- Experience with configuration management and automation tools such as Ansible, Puppet, or Chef
- Experience operating monitoring, observability, and incident response platforms such as PagerDuty, Grafana, Prometheus, ELK/OpenSearch, New Relic, or Sensu
- Track record of practical problem solving, RCA, and corrective actions that improve service reliability
- Strong understanding of networking, the OSI model, SQL/NoSQL databases, and troubleshooting distributed systems
- Ability to prioritize tasks, work independently, and communicate clearly with technical and non-technical audiences
- Enthusiasm for collaborating with globally distributed teams through video conferencing, Slack, and other communication tools
Desired Qualifications
- Experience contributing to large-scale AKS migrations, container platform modernization, or cloud transformation programs
- Proven ability to coordinate complex production releases across Engineering, Operations, and stakeholder teams
- Exposure to AI-assisted operations, generative AI tools, or automation use cases that improve engineering productivity
- Experience building reliability tooling, self-healing capabilities, or automation frameworks that reduce operational toil
- Relevant certifications in Azure, AWS, Terraform, container platforms, or cloud reliability engineering
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.