O.C. Tanner logo
O.C. TannerPosted 1 week ago

Manager, Site Reliability Engineering

On-siteSalt Lake City, Utah, United States

Full TimeLarge

Job Summary

Lead a team of Site Reliability Engineers to define and execute reliability strategy, improving availability and resilience through automation and observability. Establish priorities, goals, and metrics aligned with business objectives, then partner with Engineering, Product, and Support to drive shared ownership of production services. Build enterprise standards for metrics, logs, and traces using OpenTelemetry, Datadog, and Coralogix while overseeing production triage, incident response, and blameless post-mortems. Manage on-call programs, workforce planning, and capacity to ensure seamless 24x7 coverage in a follow-the-sun model. Report reliability trends and risks to executive leadership.

Required Qualifications

  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines, including 2+ years in a technical leadership or people management role
  • Proven experience leading teams responsible for production operations, reliability engineering, incident management, and operational excellence
  • Experience designing and implementing SRE practices, reliability programs, or operational maturity initiatives within growing engineering organizations
  • Experience operating large-scale, customer-facing SaaS platforms with high availability, performance, and scalability requirements
  • Strong understanding of modern software engineering practices and partnering with development teams to build reliable, resilient systems
  • Hands-on experience with observability platforms such as OpenTelemetry, Datadog, Coralogix, or similar technologies
  • Strong knowledge of AWS and Kubernetes in production environments
  • Deep understanding of monitoring, logging, distributed tracing, SLIs, SLOs, error budgets, and reliability engineering principles
  • Demonstrated ability to lead cross-functional initiatives and influence stakeholders across Engineering, Product, and Support organizations
  • Experience developing engineering roadmaps, defining team objectives, aligning reliability investments with business priorities, and driving continuous operational improvement through incident learning and post-incident reviews

Desired Qualifications

  • Experience leading distributed or globally dispersed engineering teams
  • Experience with multiple cloud providers or cloud-agnostic platform architectures
  • Familiarity with security, compliance, governance, and operational risk management frameworks
  • Proficiency with modern Infrastructure-as-Code and technologies such as Terraform, Golang, Python, Playwright, and Performance Monitoring tools
  • Experience with relational and distributed data technologies such as PostgreSQL, OpenSearch, Redis/ElastiCache, or Aurora
  • Experience with messaging and streaming platforms such as Kafka, ActiveMQ, SNS/SQS, or similar event-driven technologies
  • Strong understanding of cost optimization, platform sustainability, and engineering efficiency metrics

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce