Site Reliability Engineer
$55,000–$75,000 year
On-siteDublin, Leinster, Ireland
Job Summary
Conduct daily operations with a hyper focus on triage, root cause analysis, and blameless post-mortems to understand the business impact of products. Partner with developers to design, build, and support technology services, ensuring operational criteria like system availability, capacity, and performance are met. Automate data-driven alerts to proactively escalate issues and establish SLOs to improve reliability. Manage the CI/CD pipeline for promoting software into higher environments through validation and operational gating. This Business Operations Site Reliability Engineer role leads the DevOps transformation at Client, fostering developer run ownership and empowering teams to build resilient products. The position requires immediate to one-month notice and offers a salary up to 75k for a permanent, Dublin-based hybrid role.
Required Qualifications
- BS degree in Computer Science or related technical field involving coding (e.g., physics or mathematics), or equivalent practical experience
- Coding and/or scripting exposure
- Appetite for change and pushing the boundaries of what can be done with automation
- Experience with algorithms, data structures, scripting, pipeline management, and software design
- Systematic problem-solving approach, coupled with strong communication skills and a sense of ownership and drive
- Interest in designing, analysing, and troubleshooting large-scale distributed systems
- Willingness and ability to learn and take on challenging opportunities and to work as a member of matrix based diverse and geographically distributed project team
- Ability to balance doing things right with fixing things quickly
- Comfortable collaborating with cross-functional teams to ensure that expected system behavior is understood, and monitoring exists to detect anomalies
- Kafka Knowledge
- 3-5 years of experience working with Apache Kafka in a production environment
- Strong knowledge of Kafka architecture, including brokers, topics, partitions, and replicas
- Experience with Kafka security, including SSL, SASL, and ACLs
- Proficiency in configuring, deploying, and managing Kafka clusters in cloud and on-premises environments
- Experience with Kafka stream processing using tools like Kafka Streams, KSQL, or Apache Flink
- Solid understanding of distributed systems, data streaming, and messaging patterns
- Proficiency in Java, Scala, or Python for Kafka-related development tasks
- Familiarity with DevOps practices, including CI/CD pipelines, monitoring, and logging
- Experience with tools like Zookeeper, Schema Registry, and Kafka Connect
- Strong problem-solving skills and the ability to troubleshoot complex issues in a distributed environment
- Excellent communication and collaboration skills to work effectively with cross-functional teams and stakeholders
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.