Lead SRE - Chase UK
On-siteLondon, England, United Kingdom
Job Summary
Drive continuous improvement of reliability, monitoring, and alerting for mission-critical microservices while reducing operational toil through automation and infrastructure tooling. Develop meaningful service metrics, error budgets, and dashboards to proactively identify bottlenecks via performance testing and capacity planning. Design self-healing patterns, resiliency strategies, and failover approaches, partnering with engineering, product, and platform teams to embed reliability standards throughout the software lifecycle. Leverage approved AI tools to accelerate root-cause analysis, runbook drafting, and post-incident investigations, establishing governance for AI-assisted engineering practices across the team.
Required Qualifications
- Formal training or certification on software engineering concepts
- Advanced applied experience in software engineering
- Proven experience as a software engineer
- Proficiency in at least one programming language such as Python, Go, or Java
- Demonstrated experience designing, coding, testing, and delivering software in at least one technology stack
- Strong debugging and troubleshooting skills across distributed systems
- Demonstrated experience as a Site Reliability Engineer or Site Reliability Engineer supporting production services
- Working knowledge of microservice infrastructure components, including service discovery, ingress, networking, and load balancing
- Experience with Kubernetes
- Experience with cloud computing services
- Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger
- Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive information
- Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment
- Ability to set team expectations for validating AI outputs for correctness, performance, and security
- Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations
- Experience coaching senior engineers/leads on compliant usage patterns and controls
Desired Qualifications
- Experience with AWS
- Experience building internal tooling for reliability, including command-line tools, automation pipelines, operators, or controllers
- Experience improving developer experience through golden paths, paved roads, templates, or reusable engineering patterns
- Experience applying AI to operational workflows such as alert enrichment, summarisation, runbook generation, or anomaly triage using approved tools and patterns
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.