NextGen Healthcare logo
NextGen HealthcarePosted 1 month ago

Sr. Cloud Operations Reliability Engineer (SRE)

RemoteUnited States

Full TimeSenior LevelBachelors DegreeLargeHealthcare Tech

Job Summary

Drive operational excellence by owning critical reliability initiatives, establishing observability and service health practices, and leading incident response coordination. Design and implement reliability-focused automation, operational tooling, and runbooks to reduce manual toil, while building comprehensive monitoring, logging, and alerting strategies. Conduct performance and capacity analysis to identify bottlenecks and provide recommendations for scaling and operational readiness. Partner with engineering teams to evaluate deployment readiness, support disaster recovery planning, and mentor team members on reliability standards. Establish and maintain SLOs/SLIs, apply Infrastructure as Code practices, and ensure compliance with cloud governance and security initiatives.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field
  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline
  • demonstrated ownership of production systems
  • Extensive hands-on experience supporting production cloud environments using Google Cloud Platform (GCP), AWS, or equivalent cloud service providers
  • Proven expertise in monitoring, observability platforms, alerting strategies, incident response, root cause analysis, and production support in distributed or cloud-native architectures
  • Demonstrated experience with Infrastructure as Code (Terraform, Deployment Manager, CloudFormation, etc.)
  • version control best practices
  • experience driving corrective actions and establishing reliability improvements
  • Experience with Kubernetes operations, containerization, and orchestration platforms
  • Experience with application performance monitoring (APM) and distributed tracing
  • experience mentoring junior engineers or leading operational improvements initiatives
  • Google Cloud certifications: Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer
  • AWS certification: AWS SysOps Administrator or equivalent
  • Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering (SRE), or ITIL
  • Working knowledge of CI/CD practices
  • cloud governance
  • compliance frameworks
  • disaster recovery
  • business continuity planning
  • Deep technical knowledge of Google Cloud Platform (GCP), AWS, or similar cloud providers
  • understanding of cloud-native services, networking, security, and compute models
  • Hands-on expertise with monitoring platforms (Datadog, New Relic, Prometheus, Cloud Monitoring, etc.)
  • ability to design effective dashboards, alerts, and health checks
  • Proficiency in scripting languages (Python, Bash, Go, etc.)
  • Advanced ability to diagnose complex, multi-layered infrastructure issues and coordinate timely recovery
  • Ability to translate complex technical findings into actionable recommendations
  • experience influencing cross-functional teams on reliability practices

Desired Qualifications

  • Security operations
  • compliance auditing
  • audit readiness processes

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce