Site Reliability Engineer - Observability
$35,000–$90,000 year
On-siteGalway, Connacht, Ireland
Job Summary
Drive adoption of CI/CD pipelines, Infrastructure as Code, and GitOps practices to standardize deployment workflows and champion developer productivity. Lead the evolution from traditional monitoring to AI-driven operations by integrating machine learning models for predictive monitoring, anomaly detection, and automated root cause analysis. Design scalable observability frameworks covering metrics, logs, and traces to enable real-time insights across microservices and cloud-native architectures. Participate in incident triage, implement runbooks and automated escalations, and establish metrics-driven approaches to measure system health and availability.
Required Qualifications
- 1+ years of experience in software engineering, SRE, or production engineering in large-scale distributed systems
- Hands-on experience with Observability tools such as AppDynamics, Grafana, Prometheus, Datadog, OpenTelemetry, or similar
- Experience with AIOps or intelligent monitoring platforms, including anomaly detection and event correlation
- Strong expertise in cloud platforms (AWS, Azure, or GCP), cloud-native architectures (Kubernetes, containers, microservices), and CI/CD pipelines (GitHub Actions, Jenkins)
- Proficiency in at least one programming language (e.g., Python, Java, Go)
- Strong understanding of distributed systems, resiliency patterns, and fault tolerance
- Experience implementing incident management, on-call processes, and root cause analysis
- Hands-on expertise with Infrastructure as Code (Terraform, ARM, CloudFormation) and CI/CD pipelines
- Experience using GenAI/Automation tools and frameworks such as OpenAI, CoPilot, Gemini, Claude, MCP etc
- Proven ability to design scalable, reliable, and observable systems
Desired Qualifications
- Strong knowledge of machine learning applications in IT operations (e.g., anomaly detection, forecasting, clustering)
- Experience defining and managing SLIs/SLOs and error budgets at scale
- Experience with OpenTelemetry and modern observability standards
- Familiarity with chaos engineering, resilience testing, and fault injection frameworks
- Exposure to GenAI-driven operations or AI-assisted troubleshooting tools
- Demonstrated leadership in driving cross-functional initiatives and influencing senior stakeholders
- Contributions to open-source projects in SRE, observability, or AIOps domains
- Certifications in AIOps, SRE, OpenTelemetry, cloud platforms, or DevOps
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.