Lead Site Reliability Engineer
On-siteDublin, Leinster, Ireland
Job Summary
Engineer reliability into enterprise-scale platforms by writing production software for automation, control loops, and self-healing systems that remove manual operations. Build declarative, intent-based designs using trusted sources of truth to reconcile reality, while treating telemetry as a first-class asset for detection and closed-loop remediation. Define SLIs and SLOs with stakeholders, implement SLO-based alerting, and own services end-to-end for reliability, performance, and cost. Share an on-call rotation, lead incident triage and post-mortems, and drive down toil through measurable automation. Work AI-native across the SDLC with strict validation standards to accelerate development without compromising correctness. Set reliability standards across the team and partner organizations while applying security-first judgment throughout the engineering lifecycle.
Required Qualifications
- Solid production coding experience in an industry-standard language (e.g. Python, Go, Java, C++, Rust)
- Experience running production systems at scale, including on-call ownership, incident response, and designing for reliability and operability
- Experience with SLI/SLO/error-budget practice, or clear aptitude and appetite to own it
- Observability depth: white-box/black-box monitoring, SLO-based alerting, and telemetry, using tools such as Grafana, Prometheus, Splunk, Datadog, Dynatrace, or equivalent
- Experience with *nix and with infrastructure automation and tooling (e.g. Kubernetes, Terraform, CI/CD)
- Strong systems thinking: interfaces, contracts, failure modes, and interactions at scale
- Fluency directing AI tools to do real engineering work, not just autocomplete, with sound judgment on where AI applies and where deep human expertise is required
- Security-first mindset, integrating risk judgment from design through production
- Clear, direct communication and calm ownership under pressure during high-severity events
- Outcome orientation: focused on reliability, impact, and cost, not activity
Desired Qualifications
- Networking depth (routing, switching, security, packet/flow analysis) or experience operating network-adjacent platforms
- Experience across multiple infrastructure domains or programming languages
- Demonstrated ongoing AI skill development (e.g. context/prompt engineering, agent orchestration) and use of AI to redesign workflows for measurable impact
- Prior experience in regulated or large-scale enterprise environments
- Experience establishing engineering culture
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.