Head of Site Reliability Engineering (SRE)
On-siteLondon, England, United Kingdom
Job Summary
Establish a multi-year platform and reliability strategy, defining the SRE roadmap and driving engineering standards for observability, automation, and production readiness. Lead, mentor, and scale a global team of senior SREs while architecting a secure hybrid infrastructure across GCP, AWS, and on-premise data centers using Kubernetes and Hashicorp Nomad. Own incident response and root-cause analysis to meet MTTD/MTTR targets, govern cloud architecture and cost management, and build relationships with departments to deliver services while maintaining auditor and regulator standing. This role supports the company's growth in global crypto financial services, requiring a principal leader to enhance infrastructure vision and ensure uptime and safety for over 90 million wallet holders.
Required Qualifications
- Proven experience leading senior SRE teams in 24/7 financial environments
- Ability to influence alignment around shared engineering standards
- Deep expertise in GCP and AWS networking/management
- Hands-on experience with Kubernetes
- Hands-on experience with Nomad
- Optimized Terraform workflows
- Solid background in bare-metal hardware
- Data center networking
- Implementing Cloudflare Zero Trust architectures
- Strong software engineering background
- Deep understanding of distributed systems
- Deep understanding of Linux
- Approaches to achieving good observability & reliability
- Strong judgment in prioritizing reliability, security, and developer productivity
- Evolving complex, heterogeneous estates
- Mandatory in-office presence four days per week
Desired Qualifications
- Experience with Hashicorp Nomad
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.