Senior Software Engineer - DevOps, SRE, Platform Engineering
$50,000–$125,000 year
On-siteGalway, Connacht, Ireland
Job Summary
Drive CI/CD pipeline adoption, Infrastructure as Code, and GitOps practices to standardize automated, secure platform engineering solutions. Lead the evolution from traditional monitoring to AI-driven operations by integrating machine learning for predictive monitoring, anomaly detection, and automated root cause analysis. Architect scalable observability frameworks covering metrics, logs, and traces while defining instrumentation standards for cloud-native architectures. Manage enterprise-wide SRE practices including SLIs, SLOs, and error budgets to establish metrics-driven approaches for system health. Lead incident response, RCA, and postmortems to build self-healing systems that reduce Mean Time to Resolution. This role advances intelligent, automated reliability practices across the CVS Digital team, influencing architecture and mentoring engineers to embed observability into the software delivery lifecycle.
Required Qualifications
- 3+ years of experience in software engineering, SRE, or production engineering in large-scale distributed systems
- Hands-on experience with Observability tools such as AppDynamics, Grafana, Prometheus, Datadog, OpenTelemetry, or similar
- Experience with AIOps or intelligent monitoring platforms, including anomaly detection and event correlation
- Strong expertise in cloud platforms (AWS, Azure, or GCP), cloud-native architectures (Kubernetes, containers, microservices), and CI/CD pipelines (GitHub Actions, Jenkins)
- Proficiency in at least one programming language (e.g., Python, Java, Go)
- Strong understanding of distributed systems, resiliency patterns, and fault tolerance
- Experience implementing incident management, on-call processes, and root cause analysis
- Hands-on expertise with Infrastructure as Code (Terraform, ARM, CloudFormation) and CI/CD pipelines
- Experience using GenAI/Automation tools and frameworks such as OpenAI, CoPilot, Gemini, Claude, MCP etc
- Proven ability to design scalable, reliable, and observable systems
Desired Qualifications
- Experience designing and implementing AIOps platforms or predictive reliability systems at scale
- Strong knowledge of machine learning applications in IT operations (e.g., anomaly detection, forecasting, clustering)
- Experience defining and managing SLIs/SLOs and error budgets at scale
- Experience with OpenTelemetry and modern observability standards
- Familiarity with chaos engineering, resilience testing, and fault injection frameworks
- Exposure to GenAI-driven operations or AI-assisted troubleshooting tools
- Experience in healthcare, finance, enterprise SaaS, or highly regulated industries
- Demonstrated leadership in driving cross-functional initiatives and influencing senior stakeholders
- Contributions to open-source projects in SRE, observability, or AIOps domains
- Certifications in AIOps, SRE, OpenTelemetry, cloud platforms, or DevOps
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.