ML Ops Lead
On-siteLondon, England, United Kingdom
Job Summary
Lead and manage a dedicated DevOps team while architecting, maintaining, and optimizing robust MLOps/LLMOps pipelines and CI/CD frameworks for continuous model deployment. Define the infrastructure roadmap for AI/ML workloads using Infrastructure as Code and deploy Large Language Models into production environments ensuring high availability and optimal performance. Establish FinOps frameworks to track and forecast AI infrastructure spend, implementing auto-scaling and down-scaling policies to eliminate waste and manage high-cost GPU/CPU budgets. Set up 24/7 incident response, telemetry, and observability metrics to monitor system performance, model drift, and data pipelines while enforcing strict data governance and security protocols across all AI/ML infrastructure.
Required Qualifications
- Extensive production experience deploying and supporting ML systems
- Proven track record of leading engineering teams
- Demonstrated experience with Generative AI and LLM deployment patterns
- A proven history of reducing cloud spend on large-scale AI clusters
- Experience with tools like MLflow, Kubeflow, LangSmith, or Phoenix
- Expertise in AWS/GCP/Azure cost tools, Kubecost, or Cloudability
- Extensive background of Kubernetes (K8s), Docker, and service meshes
- Expert knowledge of Terraform, Ansible, Jenkins, or GitHub Actions
- Proficient in Python, Bash, or Go
- Familiarity with Triton Inference Server, vLLM, or Hugging Face TGI
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.