Site Reliability Engineer
RemoteLondon, England, United Kingdom or Los Angeles, California, United States
Job Summary
Site Reliability Engineer at Epidemic Sound responsible for building and operating the platform that services rely on, including GKE-based clusters, CI/CD pipelines, traffic management, and IAM governance. You will implement and operate infrastructure-defined platforms, improve production readiness, and drive observability through metrics, tracing, runbooks, and SLOs while collaborating with product teams to enable safe, scalable releases. Familiarity with containerized platforms, cloud-native tooling, and an agentic approach to development (including AI-assisted workflows) is highly valued; optional exposure to GCP, Prometheus, Thanos, Grafana, and service meshes like Linkerd, Istio, or Cilium is a plus. Remote work is available. Equal opportunity employer.
Required Qualifications
- Kubernetes fundamentals: a solid grasp of controllers, core components, and CNI and networking - depth in the domain matters more than any single tool (GKE a plus).
- Infrastructure as code and delivery: Terraform, Helm or Kustomize, CI/CD and GitOps (ArgoCD), and the traffic-management and progressive-delivery mechanisms that move releases out safely.
- Networking and access: routing fundamentals, the VPC, firewall, and network-policy primitives beneath it, and IAM and access management at different levels.
- Operational depth: monitoring fundamentals (a clear view of when to reach for metrics versus tracing, and experience with an open-source observability stack), strong troubleshooting across distributed systems, and solid Unix/Linux.
- Agentic development mindset: you use AI agents actively in your own work, knowing where they add leverage and where human judgement is non-negotiable.
- Collaboration and judgement: you do your best work on large, cross-cutting projects, communicate openly, and stay opinionated but open to discussion - reaching for the right tool over your own creation.
- It would also be music to our ears if you have
- Familiarity with GCP and an observability stack with Prometheus, Thanos, and Grafana.
- Experience running containerised platforms at scale.
- Service mesh experience with Cilium eBPF, Linkerd, or Istio.
- Familiarity with platform building blocks like cert-manager, external-secrets, or external-dns.
Desired Qualifications
- Kubernetes fundamentals
- Infrastructure as code and delivery (Terraform, Helm, Kustomize)
- CI/CD and GitOps (ArgoCD)
- Networking and access management (VPCs, firewalls, IAM)
- Operational depth with monitoring and observability stacks
- Experience with UNIX/Linux environments
- Ability to collaborate on large cross-cutting projects
- Familiarity with GCP and observability tools (Prometheus, Grafana, Thanos)
- Experience with container platforms at scale
- Service mesh experience (e.g., Cilium, Linkerd, Istio)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.