Sr. Cloud Operations Reliability Engineer (SRE)
RemoteUnited States
Job Summary
Drive operational excellence by owning critical reliability initiatives, establishing observability and service health practices, and leading incident response coordination. Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, while building comprehensive monitoring, logging, and alerting strategies. Conduct performance and capacity analysis to identify bottlenecks and provide recommendations for scaling and operational readiness. Partner with engineering teams to evaluate deployment readiness, support disaster recovery planning, and mentor team members on reliability standards. Establish and maintain SLOs/SLIs, apply Infrastructure as Code practices, and ensure compliance with cloud governance and security initiatives.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field
- 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline
- demonstrated ownership of production systems
- Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers
- Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures
- Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.)
- version control best practices
- experience driving corrective actions and establishing reliability improvements
- Experience with Kubernetes operations, containerization, and orchestration platforms
- Experience with application performance monitoring (APM) and distributed tracing
- experience mentoring junior engineers or leading operational improvements initiatives
- Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer
- AWS certification: AWS SysOps Administrator or equivalent
- Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL
- Working knowledge of CI/CD practices
- cloud governance
- compliance frameworks
- disaster recovery
- business continuity planning
- Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers
- understanding of cloud-native services, networking, security, and compute models
- Hands-on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.)
- ability to design effective dashboards, alerts, and health checks
- Proficiency in scripting languages (Python, Bash, Go, etc.)
- Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery
- Ability to translate complex technical findings into actionable recommendations
- experience influencing cross-functional teams on reliability practices
Desired Qualifications
- Security operations
- compliance auditing
- audit readiness processes
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.