SDM SRE Senior Engineer
On-siteHyderabad, Telangana, India
Job Summary
Implement SRE best practices across production support, incident management, and release processes while driving shift-left reliability in Azure Cloud. Build automation scripts, self-healing workflows, and operational tools to reduce manual toil and improve mean time to resolution. Configure and enhance observability across logs, metrics, and traces using tools like Splunk, Datadog, and Azure Monitor to support SLIs, SLOs, and operational dashboards. Contribute to AI-driven operational capabilities including incident summarization, anomaly detection, and event correlation to accelerate incident response. Collaborate with application, infrastructure, cloud, security, and leadership teams to strengthen reliability and modernize production service support.
Required Qualifications
- Bachelor's degree in Information Technology, Computer Science, Engineering, or equivalent practical experience
- 6+ years of experience in SRE, DevOps, Production Support Engineering, Cloud Operations, Automation Engineering, or related roles
- Hands-on experience in application support, incident management, problem management, change/release management, deployment support, monitoring, and documentation
- Experience supporting production applications and platforms in large-scale enterprise environments
- Experience working across application, infrastructure, cloud, operations, security, and leadership teams
- Experience working within ITIL-based operational environments
- Strong experience with observability/APM tools such as Splunk, AppDynamics, Dynatrace, Datadog, New Relic, Azure Monitor, or similar tools
- Experience with SLIs, SLOs, SLAs, operational KPIs, service health dashboards, and reliability metrics
- Strong coding/scripting experience in one or more languages such as Python, Java, PowerShell, Go, or Shell scripting
- Hands-on experience with automation and DevOps tools such as Azure DevOps, GitHub Actions, Jenkins, Ansible, Terraform, Power Automate, Rundeck, or similar tools
- Knowledge of AIOps, Gen AI, or AI-enabled operations use cases such as anomaly detection, log analysis, ticket classification, event correlation, and incident summarization
- Experience integrating enterprise tools using APIs, webhooks, scripts, and automation workflows
- Hands-on experience supporting or implementing solutions in Azure Cloud
- Understanding of distributed systems, APIs, microservices, cloud-native applications, and reliability engineering principles
- Strong troubleshooting, analytical, communication, and stakeholder management skills
- Experience with Agentic AI, AI agents, or Gen AI-based automation for IT operations
- Experience with Azure OpenAI, Microsoft Copilot Studio, LangChain, Semantic Kernel, vector databases, or RAG-based solutions
- Experience with ITSM and incident response tools such as ServiceNow, Jira Service Management, PagerDuty, xMatters, Opsgenie, or similar platforms
- Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code, GitOps, or DevSecOps
- Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time-series analysis
- Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls
- Experience in the retail domain
- Relevant certifications in Azure, DevOps, SRE, AI, or ITIL
- Strong analytical and problem-solving mindset
- Automation-first and continuous improvement mindset
- Strong communication and stakeholder management skills
- Ability to influence without direct authority
- Ability to work across product, engineering, operations, cloud, security, and leadership teams
- Curiosity and willingness to explore AIOps, Gen AI, and Agentic AI capabilities
- Ability to balance innovation with reliability, stability, security, and compliance
- Customer-focused approach with strong ownership and accountability
- Comfortable supporting flexible working hours as needed to engage with global stakeholders, participate in critical meetings, support escalations, and ensure seamless delivery across regions
Desired Qualifications
- Experience with Agentic AI, AI agents, or Gen AI-based automation for IT operations
- Experience with Azure OpenAI, Microsoft Copilot Studio, LangChain, Semantic Kernel, vector databases, or RAG-based solutions
- Experience with ITSM and incident response tools such as ServiceNow, Jira Service Management, PagerDuty, xMatters, Opsgenie, or similar platforms
- Experience with Docker, Kubernetes, OpenShift, CI/CD pipelines, Infrastructure as Code, GitOps, or DevSecOps
- Knowledge of machine learning concepts such as anomaly detection, classification, clustering, and time-series analysis
- Understanding of Responsible AI, prompt engineering, model governance, data privacy, and security controls
- Experience in the retail domain
- Relevant certifications in Azure, DevOps, SRE, AI, or ITIL
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.