Site Reliability Engineer III (DSS)
On-siteDublin, Leinster, Ireland
Job Summary
Guide peers in adopting site reliability engineering best practices and design appropriate level designs. Collaborate with software engineers to implement automated continuous integration and delivery pipelines while configuring infrastructure, configuration, and network as code. Utilize enterprise-authorized AI capabilities to accelerate incident triage, troubleshoot issues, and identify operational patterns indicating reliability risks. Resolve complex problems using service level indicators and objectives, leveraging observability tools like Grafana, Dynatrace, and Prometheus to monitor applications and platforms. Work with technical experts to improve availability, reliability, and scalability, prioritizing reuse-first improvements tied to SLO outcomes.
Required Qualifications
- Formal training or certification on site reliability engineering concepts
- Advanced applied experience in site reliability engineering
- Proficiency in site reliability culture and principles
- Familiarity with how to implement site reliability within an application or platform
- Proficiency in at least one programming language such as Python, Java/Spring Boot, and .Net
- Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
- Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements, while ensuring compliance with risk controls and company-wide standards
- Proficient knowledge of software applications and technical processes within a given technical discipline (e.g., Cloud, AI & etc.)
- Hands-on experience supporting cloud and data platforms, including AWS (e.g., EC2, Athena) and Snowflake, with a reliability-first mindset
- Experience in observability such as white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, AppDynamics, and others
- Experience with continuous integration and continuous delivery tooling (e.g., Jenkins, GitLab, Terraform)
- Familiarity with container and container orchestration and troubleshooting common networking technologies and issues (e.g., Docker, Kubernetes, ECS)
Desired Qualifications
- Strong understanding of release engineering practices, GitOps, branching strategies, and change management discipline
- Experience with incident management, on-call operations, capacity planning, and reliability engineering practices
- Experience authoring pipeline configuration (for example, YAML) and infrastructure templates (for example, Terraform or ARM)
- AWS and Azure experience
- Oracle Cloud Infrastructure (OCI) exposure
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.