Boehringer Ingelheim logo
Boehringer IngelheimPosted 2 months ago

Senior MLOps Engineer

HybridLondon, England, United Kingdom

Full TimeSenior LevelMasters DegreeEnterprise

Job Summary

Senior MLOps Engineer to join AI Enablement in London, focused on moving production ML models from development to production with strong ownership of deployment, monitoring, retraining and lifecycle. Responsibilities include ensuring experiment tracking and model registry provenance, configuring distributed training and fine-tuning jobs, participating in model handovers with ML engineers, deploying and monitoring model serving endpoints, and upholding MLOps standards across the accelerator. Role is hybrid with ~3 days in the office and requires MSc in a related technical field (PhD preferred) and hands-on production ML experience; familiarity with PyTorch Distributed/DeepSpeed/Ray Train, MLflow/W&B, CI/CD for ML, cloud infrastructure, Terraform, and biomedical AI workloads.

Required Qualifications

  • MSc in Machine Learning, Computer Science, Software Engineering or a related technical field; PhD preferred or the equivalent industry experience
  • Solid hands-on experience operating ML training and serving workflows in production environments
  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, FSDP or Ray Train
  • Experience operating experiment tracking systems and model registry systems such as MLflow, Weights and Biases or equivalent
  • Familiarity with CI/CD tooling for ML workflows e.g. cloud-native pipeline services, GitHub Actions or equivalent
  • Solid understanding of cloud infrastructure for ML (compute, storage, networking) that is sufficient to specify requirements clearly and diagnose infrastructure-related issues
  • Awareness of large model training characteristics including memory footprint, compute scaling and parallelisation strategies
  • Familiarity with infrastructure-as-code tooling such as Terraform or cloud-native equivalents
  • Experience working closely with research and ML engineering teams as a platform operator
  • Experience operating ML infrastructure for large foundation model training
  • Familiarity with biomedical AI workloads, such as training foundation models on large-scale multimodal data

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce