Barracuda Networks logo
Barracuda NetworksPosted 1 week ago

Manager, Cloud Services and Site Reliability

$150,000–$200,000 year

HybridOttawa, Ontario, Canada

Full TimeLarge

Job Summary

Lead and develop a high-performing Site Reliability Engineering team, setting clear expectations and fostering a culture of ownership for critical SaaS applications. Drive reliability practices including SLOs, SLIs, monitoring, alerting, and capacity planning while partnering with engineering and product teams to improve system design and scalability. Own incident management, coordinate major incidents, and lead post-incident reviews to implement systemic reliability improvements. Champion automation and tooling to reduce operational toil and scale support for production services, using operational data to prioritize reliability gaps and communicate progress to stakeholders. Support secure and compliant operations by embedding appropriate controls and documentation into service delivery.

Required Qualifications

  • 5+ years of experience in SRE, DevOps, infrastructure, cloud operations, or a related technical operations discipline, including experience leading or managing technical teams
  • Strong understanding of cloud platforms, distributed systems, production operations, and modern reliability practices
  • Experience implementing or improving SLOs, SLIs, monitoring, alerting, incident response, and post-incident review practices
  • Demonstrated ability to hire, mentor, coach, and develop engineers while building a healthy, accountable, and inclusive team culture
  • Strong communication skills, with the ability to explain technical topics clearly to engineering partners, product stakeholders, and business leaders
  • Track record of using data, operational insight, and structured problem solving to improve service reliability and team effectiveness
  • Experience with infrastructure automation, CI/CD practices, disaster recovery, cost optimisation, or multi-cloud operations
  • Experience influencing operational change across teams, improving documentation practices, or evaluating tools and vendors that support service reliability
  • Hybrid position based in Ottawa, Ontario

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce