CVS Health logo
CVS HealthPosted 3 weeks ago

Staff Software Engineer - DevOps, SRE, AIOps

$65,000–$165,000 year

On-siteGalway, Connacht, Ireland

Full TimeSenior LevelLargeHEALTHCARE

Job Summary

Lead DevOps, SRE, and AIOps initiatives to drive adoption of CI/CD pipelines, Infrastructure as Code, and GitOps practices across the CVS Digital team. Architect scalable observability frameworks covering metrics, logs, and traces while implementing enterprise-wide SRE standards including SLIs, SLOs, and error budgets. Integrate machine learning models for predictive monitoring, anomaly detection, and automated root cause analysis to reduce operational toil and accelerate incident resolution. Build self-healing systems, automate remediation workflows, and establish on-call readiness to minimize Mean Time to Resolution. Mentor engineers and influence architecture decisions to embed reliability into the software delivery lifecycle.

Required Qualifications

  • 5+ years of experience in software engineering, SRE, or production engineering in large-scale distributed systems
  • Hands-on experience with Observability tools such as AppDynamics, Grafana, Prometheus, Datadog, OpenTelemetry, or similar
  • Experience with AIOps or intelligent monitoring platforms, including anomaly detection and event correlation
  • Strong expertise in cloud platforms (AWS, Azure, or GCP), cloud-native architectures (Kubernetes, containers, microservices), and CI/CD pipelines (GitHub Actions, Jenkins)
  • Proficiency in at least one programming language (e.g., Python, Java, Go)
  • Strong understanding of distributed systems, resiliency patterns, and fault tolerance
  • Experience implementing incident management, on-call processes, and root cause analysis
  • Hands-on expertise with Infrastructure as Code (Terraform, ARM, CloudFormation) and CI/CD pipelines
  • Experience using GenAI/Automation tools and frameworks such as OpenAI, CoPilot, Gemini, Claude, MCP etc
  • Proven ability to design scalable, reliable, and observable systems

Desired Qualifications

  • Experience designing and implementing AIOps platforms or predictive reliability systems at scale
  • Strong knowledge of machine learning applications in IT operations (e.g., anomaly detection, forecasting, clustering)
  • Experience defining and managing SLIs/SLOs and error budgets at scale
  • Experience with OpenTelemetry and modern observability standards
  • Familiarity with chaos engineering, resilience testing, and fault injection frameworks
  • Exposure to GenAI-driven operations or AI-assisted troubleshooting tools
  • Experience in healthcare, finance, enterprise SaaS, or highly regulated industries
  • Demonstrated leadership in driving cross-functional initiatives and influencing senior stakeholders
  • Contributions to open-source projects in SRE, observability, or AIOps domains
  • Certifications in AIOps, SRE, OpenTelemetry, Cloud platforms, or DevOps

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce