Barracuda Networks logo
Barracuda NetworksPosted 1 week ago

Cloud Site Reliability Senior Engineer

On-siteBengaluru, Karnataka, India

Full TimeSenior LevelBachelors DegreeLarge

Job Summary

Lead reliability improvements for production services by driving troubleshooting, incident response, and RCA across cloud services. Partner with Engineering and Operations to enhance release readiness, deployment practices, and operational standards while strengthening observability and alerting. Support cloud platform improvements, automation, and resilience initiatives using Terraform-managed infrastructure on Azure and AWS to reduce toil. Strengthen operational insights and share knowledge to reduce silos across SRE and Cloud Operations teams. Contribute to secure, compliant, and cost-conscious operations for containerized services and AKS environments.

Required Qualifications

  • Bachelor's degree in computer science engineering, Information Technology, or equivalent degree
  • 5+ years of progressive experience in Site Reliability Engineering, Cloud Operations, DevOps, or Platform Engineering
  • Strong Linux/Unix command-line administration, troubleshooting, package management, and systems operations skills
  • Experience managing production cloud infrastructure across Azure and/or AWS environments
  • Hands-on experience operating container orchestration platforms, including AKS migrations, upgrades, and production operations
  • Hands-on Infrastructure as Code experience, preferably with Terraform for provisioning and lifecycle management
  • Automation, scripting, and AI-assisted engineering experience using Python, Bash, Go, YAML, generative AI tools, or similar technologies
  • Experience with CI/CD pipelines using Azure DevOps, ArgoCD, Jenkins, GitHub Actions, or similar platforms
  • Experience with Docker, container registries such as Azure Container Registry, and container lifecycle management
  • Experience with configuration management and automation tools such as Ansible, Puppet, or Chef
  • Experience operating monitoring, observability, and incident response platforms such as PagerDuty, Grafana, Prometheus, ELK/OpenSearch, New Relic, or Sensu
  • Track record of practical problem solving, RCA, and corrective actions that improve service reliability
  • Strong understanding of networking, the OSI model, SQL/NoSQL databases, and troubleshooting distributed systems
  • Ability to prioritize tasks, work independently, and communicate clearly with technical and non-technical audiences
  • Enthusiasm for collaborating with globally distributed teams through video conferencing, Slack, and other communication tools

Desired Qualifications

  • Experience contributing to large-scale AKS migrations, container platform modernization, or cloud transformation programs
  • Proven ability to coordinate complex production releases across Engineering, Operations, and stakeholder teams
  • Exposure to AI-assisted operations, generative AI tools, or automation use cases that improve engineering productivity
  • Experience building reliability tooling, self-healing capabilities, or automation frameworks that reduce operational toil
  • Relevant certifications in Azure, AWS, Terraform, container platforms, or cloud reliability engineering

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce