Brevan Howard logo
Brevan HowardPosted 1 month ago

Senior Site Reliability Engineer

On-siteLondon, England, United Kingdom

Full TimeSenior LevelLarge

Job Summary

Architect, deploy, and maintain highly scalable and reliable infrastructure on Google Cloud Platform using Kubernetes and Infrastructure-as-Code tools. Champion automation across the software development lifecycle utilizing Python and Bash to reduce toil, while owning and evolving declarative infrastructure with Terraform and Helm. Implement robust monitoring, alerting, and logging solutions to ensure system visibility and proactive issue identification, then define and enforce Service Level Objectives to measure reliability. Work closely with software development teams to guide deployment strategies and scalability, taking full ownership of projects from inception through production operation. Participate in on-call rotation and lead post-incident reviews to drive continuous improvement.

Required Qualifications

  • 4+ years of hands-on experience with Google Cloud Platform (GCP) (or similar cloud infrastructure)
  • Expert-level proficiency in managing, scaling, and troubleshooting production Kubernetes environments
  • Deep expertise in Terraform for managing cloud and Kubernetes resources
  • Strong experience with Helm for packaging and deploying applications on Kubernetes
  • Proficient in at least one major programming language, preferably Python, for automation and tool development
  • Experience setting up and maintaining modern CI/CD pipelines
  • Practical experience implementing and managing monitoring and logging tools
  • Solid understanding of TCP/IP, load balancing, DNS, and cloud-native networking within Kubernetes
  • Strong command-line skills and experience with Linux systems
  • Demonstrated ability to own a problem end-to-end, from investigation to resolution and preventative measures
  • Happy to be a generalist and switch context quickly between support tickets, operational toil reduction, and long-term project work
  • A proven track record of rapidly learning and applying new technologies and tools
  • Excellent verbal and written communication skills for documentation and interacting with non-technical stakeholders

Desired Qualifications

  • Equivalent experience with other clouds (AWS/Azure) or similar tools
  • Familiarity with Service Mesh technologies (e.g., Istio)
  • Experience in security best practices within cloud and container environments (e.g., hardening, secrets management)
  • Certifications in GCP or Kubernetes (e.g., CKAD, CKA, Professional Cloud DevOps Engineer)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce