DeepL logo
DeepLPosted 1 month ago

Senior Research Scientist | Model Scaling

HybridLondon, England, United Kingdom

Full TimeSenior LevelDoctorate Or Professional DegreeMedium

Job Summary

Drive the selection and evaluation of open foundation models as the basis for next-generation translation systems. Lead architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and efficient designs. Design multi-capability adaptation strategies using LoRA and PEFT methods. Own the full modelling lifecycle from prototyping and ablation experiments to rigorous evaluation and production delivery. Partner with post-training and RL specialists to integrate alignment work into base models. Stay ahead of open-model literature to provide well-founded recommendations.

Required Qualifications

  • Strong hands-on experience adapting and scaling large language models via fine-tuning, instruction-tuning, or post-training of multi-billion-parameter models beyond black-box use
  • Sound judgment about architecture trade-offs at scale (e.g. dense vs. MoE) and about which open-weight foundation models to build on
  • Working knowledge of parameter-efficient and multi-capability adaptation (LoRA/PEFT and variants)
  • A hands-on builder who enjoys training models, running experiments, and debugging pipelines, and who can carry research results through to production with engineering
  • Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow)
  • Ability to communicate clearly, collaborate across teams, and align research work with product and engineering priorities
  • Drive the selection and evaluation of open foundation / open-weight models as the basis for our next-generation translation systems
  • Lead model selection and general architecture decisions for scaling to hundreds of billions of parameters, including Mixture-of-Experts and other sparse or efficient designs
  • Design multi-capability adaptation strategies using LoRA, PEFT, and related methods
  • Own the modelling lifecycle for your work: prototyping, ablations, scaling experiments, evaluation, and delivery into production, with rigorous and reproducible evaluation
  • Partner closely with post-training, RL/RLHF, and instruction-following specialists to integrate alignment and capability work into the base model
  • Stay ahead of the open-model and scaling literature, and bring well-founded recommendations back to the team

Desired Qualifications

  • Experience quantifying uncertainty in large models — calibration and confidence estimation via Bayesian methods, ensembling, steering, or prompt-based approaches
  • Experience with machine translation, multilingual NLP, or document-/layout-aware modelling experience
  • Familiarity with MoE-specific training and adaptation (e.g. expert routing, Mixture-of-LoRA-Experts) and large-scale data-mixture design

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce