Senior Cloud Engineer
On-siteLeeds, England, United Kingdom
Job Summary
Design and maintain Kubernetes infrastructure for AI workloads, including model serving, agent orchestration, and batch inference using Terraform. Build and operate CI/CD pipelines tailored to AI and agentic workflows with automated evaluation and regression checks. Define and track SLOs/SLAs, lead incident response, root cause analysis, and postmortems while participating in on-call rotations. Establish observability for AI-specific concerns such as latency, token usage, and model error rates using dashboards and alerting. Partner with the Inference Control Plane, Evals/Observability, and Context & Knowledge Platform teams to ensure operability from day one. This role requires 4+ years of DevOps/SRE experience, production Kubernetes ownership, and strong Python/Terraform skills. Join CreateFuture, an AI-native consulting partner working with clients like PayPal and adidas, offering flexible remote/hybrid options and 35 days of leave.
Required Qualifications
- 4+ years' experience in DevOps, SRE, or platform engineering
- production ownership of Kubernetes-based systems
- hands-on experience operating AI or ML systems in production
- strong in Python for automation and operational tooling
- production experience in Terraform
- experience with at least one major cloud provider (AWS or GCP)
- built and maintained CI/CD pipelines
- comfortable with observability stacks
Desired Qualifications
- experience with model serving
- experience with LLM inference
- experience with MLOps pipelines
- calm, rigorous incident management instincts
- familiarity with LLM/agent ecosystems
- familiarity with model APIs
- familiarity with vector databases
- familiarity with MCP
- familiarity with orchestration frameworks
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.