Senior Site Reliability Engineer - CloudVision
HybridDublin, Leinster, Ireland
Job Summary
Design, build, and deploy production systems with a focus on scalability, reliability, observability, and security. Develop automation solutions to eliminate toil and streamline operations across production environments. Proactively monitor systems, establish alerting strategies, and implement automated incident response to minimize downtime. Create detailed runbooks and conduct postmortem analyses to identify root causes. Collaborate with software teams to resolve infrastructural bottlenecks and optimize monitoring infrastructure using industry-standard tools. Plan maintenance windows with minimal service disruption and triage platform issues with analytical rigour. Deploy updates in staged, risk-managed rollouts while surveying best practices for secure, fault-tolerant systems. This permanent, remote role from Ireland requires 5+ years of experience and involves working with cross-functional teams to enhance product deployment workflows.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, or equivalent professional experience (5+ years in a related infrastructure or systems role)
- Proficiency in one or more programming languages: Go, Python, or bash shell scripting, with the ability to implement medium-complexity automation workflows
- Strong knowledge of Linux or UNIX from both administration and debugging perspectives
- Hands-on experience operating software systems, infrastructure, and complex applications at scale in production environments
- Demonstrated expertise in infrastructure-as-code principles and practices
- Strong problem-solving and software troubleshooting skills with a methodical, analytical approach
- Experience with server provisioning, particularly from storage and networking perspectives
- Proven ability to work collaboratively within cross-functional teams and communicate technical concepts clearly
- Experience with incident response, postmortem analysis, and continuous improvement methodologies
Desired Qualifications
- Experience with container orchestration platforms, particularly Kubernetes
- Hands-on experience with Docker and virtualisation technologies
- Proficiency in managing monitoring stacks, including Prometheus and Grafana
- Experience with CI/CD systems such as GitLab tools or Spinnaker
- Knowledge of infrastructure-as-code frameworks, particularly Terraform
- Experience managing databases such as PostgreSQL or equivalent relational database management systems
- Experience with artifact repositories and Docker registries
- Familiarity with cloud platforms (Google Cloud Platform, Amazon Web Services, or Microsoft Azure)
- Understanding of distributed systems architecture and principles
- Experience with performance tuning and system optimisation
- Knowledge of security best practices in infrastructure and systems design
- On-call support experience and comfort with incident response responsibilities
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.