Senior DevOps Engineer
On-siteLondon, England, United Kingdom
Job Summary
Manage Kubernetes platform operations including cluster lifecycle management via Cluster API, fleet-wide upgrades, bare metal provisioning, and disaster recovery planning. Operate and scale the on-prem observability stack with Mimir, Loki, Tempo, and Grafana while driving adoption across trading desks. Build reusable CI/CD components using GitLab pipelines and ArgoCD, and maintain automation for infrastructure provisioning and self-service tooling. Work directly with development teams to onboard them to the platform, write runbooks, and contribute to documentation. Join a small, high-impact team that built the entire infrastructure from scratch in 2025 to support P&L-impacting trading workloads.
Required Qualifications
- 5–8 years of experience in DevOps, SRE, or Platform Engineering roles
- Deep hands-on experience operating K8s in production
- Cluster lifecycle, troubleshooting, networking, storage, RBAC
- Experience with Cluster API, bare metal provisioning, or multi-cluster management
- Production experience with Prometheus, Grafana, and alerting
- Familiarity with Mimir, Loki, Tempo, or Thanos for scaled metrics, logs, and traces
- Understanding of OpenTelemetry (collectors, exporters, instrumentation)
- Experience with ArgoCD, Flux, or similar GitOps tools
- Building and maintaining CI/CD pipelines (GitLab CI, GitHub Actions, or equivalent)
- Artifact management (Artifactory, Nexus, or similar)
- Terraform and/or Ansible for provisioning and configuration management
- Strong fundamentals — systemd, networking, storage, performance troubleshooting
- Go or Python for automation, tooling, and scripting
- Comfortable reading and writing YAML and Helm charts
- Ability to work directly with trading desks and development teams who depend on the platforms you build
- You'll be explaining K8s concepts to people who aren't K8s experts
Desired Qualifications
- Experience working at a trading firm — understanding the urgency and reliability requirements of systems that support P&L-impacting workloads
- Experience with Kubeflow, Airflow, or ML pipeline orchestration
- Experience with advanced Kubernetes networking (CNI plugins, network policies, service mesh)
- Experience with Kafka or event streaming platforms
- Experience operating on-prem infrastructure (not just cloud) — bare metal servers, IPAM, storage systems
- Thoughtful use of AI coding assistants and interest in AI-assisted operational tooling
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.