CVS Health logo
CVS HealthPosted 1 month ago

Senior Software Engineer - DevOps, SRE, AIOps

$50,000–$125,000 year

On-siteGalway, Connacht, Ireland

Full TimeSenior LevelLargeHEALTHCARE

Job Summary

Drive adoption of CI/CD pipelines, Infrastructure as Code, and GitOps practices to standardize development and deployment workflows. Lead the evolution from traditional monitoring to AI-driven operations by integrating machine learning models for predictive monitoring, anomaly detection, and automated root cause analysis. Architect scalable observability frameworks covering metrics, logs, and traces to enable real-time insights across microservices and cloud-native architectures. Establish enterprise-wide SRE practices including SLIs, SLOs, error budgets, and reliability governance while building self-healing systems to reduce Mean Time to Resolution. Mentor engineers and influence architecture decisions to embed reliability and automation into the software delivery lifecycle.

Required Qualifications

  • 3+ years of experience in software engineering, SRE, or production engineering in large-scale distributed systems
  • Hands-on experience with Observability tools such as AppDynamics, Grafana, Prometheus, Datadog, OpenTelemetry, or similar
  • Experience with AIOps or intelligent monitoring platforms, including anomaly detection and event correlation
  • Strong expertise in cloud platforms (AWS, Azure, or GCP), cloud-native architectures (Kubernetes, containers, microservices), and CI/CD pipelines (GitHub Actions, Jenkins)
  • Proficiency in at least one programming language (e.g., Python, Java, Go)
  • Strong understanding of distributed systems, resiliency patterns, and fault tolerance
  • Experience implementing incident management, on-call processes, and root cause analysis
  • Hands-on expertise with Infrastructure as Code (Terraform, ARM, CloudFormation) and CI/CD pipelines
  • Experience using GenAI/Automation tools and frameworks such as OpenAI, CoPilot, Gemini, Claude, MCP etc
  • Proven ability to design scalable, reliable, and observable systems
  • Bachelor's degree or equivalent work experience in Computer Science, Engineering, or related discipline

Desired Qualifications

  • Experience designing and implementing AIOps platforms or predictive reliability systems at scale
  • Strong knowledge of machine learning applications in IT operations (e.g., anomaly detection, forecasting, clustering)
  • Experience defining and managing SLIs/SLOs and error budgets at scale
  • Experience with OpenTelemetry and modern observability standards
  • Familiarity with chaos engineering, resilience testing, and fault injection frameworks
  • Exposure to GenAI-driven operations or AI-assisted troubleshooting tools
  • Experience in healthcare, finance, enterprise SaaS, or highly regulated industries
  • Demonstrated leadership in driving cross-functional initiatives and influencing senior stakeholders
  • Contributions to open-source projects in SRE, observability, or AIOps domains
  • Certifications in AIOps, SRE, OpenTelemetry, cloud platforms, or DevOps

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce