Sr. Site Reliability Engineer
On-siteSalt Lake City, Utah, United States
Job Summary
Improve the availability, scalability, and performance of cloud-native applications through automation, monitoring, and engineering best practices. Build and evolve observability platforms using OpenTelemetry, Datadog, and Coralogix to establish standards for metrics, logs, and service-level objectives. Lead production triage efforts, rapidly diagnosing and resolving service disruptions while driving incident management and root cause analysis. Partner with Engineering, Support, and Product teams to embed reliability and operational excellence throughout the software development lifecycle. Participate in an on-call rotation focused on maintaining service health and automating repetitive operational tasks. Collaborate with global engineering teams in a follow-the-sun support model to ensure seamless 24x7 coverage.
Required Qualifications
- 5+ years of experience in Site Reliability Engineering, DevOps, platform engineering, or related roles
- strong background in production triage, incident response, and operational excellence
- Experience operating large-scale, customer-facing SaaS platforms with high availability and uptime requirements
- Proficiency in Go, Python, Java, or similar programming languages
- demonstrated experience building automation, production tooling, and reliability-focused engineering solutions
- Deep experience with modern Infrastructure-as-Code and GitOps technologies
- Terraform
- OpenTofu
- CDKTF
- Pulumi
- ArgoCD
- Helm
- Kubernetes
- Hands-on experience with OpenTelemetry, Datadog, Coralogix, or similar observability platforms
- Strong knowledge of AWS services
- Kubernetes in production environments
- Deep understanding of monitoring, logging, and distributed tracing for complex systems
- Ability to partner effectively with software engineering and testing teams
- design reliable systems
- improve application performance
- strengthen quality practices across the software development lifecycle
- Comfortable with participating in on-call rotations
- handling high-pressure environments
Desired Qualifications
- Experience with multiple cloud or cloud-agnostic environments
- Familiarity with security, compliance, and governance frameworks
- Experience with relational and distributed data technologies
- PostgreSQL
- OpenSearch
- Redis/ElastiCache
- Aurora
- Experience with messaging and streaming platforms
- Kafka
- ActiveMQ
- SNS/SQS
- similar event-driven technologies
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.