DeepL logo
DeepLPosted 1 month ago

Senior Research Scientist | Multimodal Systems

RemoteLondon, England, United Kingdom or United Kingdom

Full TimeSenior LevelDoctorate Or Professional DegreeMedium

Job Summary

Lead fine-tuning, post-training, and reinforcement learning for DeepL's document translation multimodal and vision models. Drive development of vision models for document, image, and media translation, ranging from ingestion to end-to-end generation. Build evaluator models for document and design quality, including rubric-based grading and reward hacking mitigation. Own the full lifecycle of model delivery, from prototyping and ablation to production deployment at scale. Establish practices for evaluation, reproducibility, and continuous model improvement. Prototype rapidly and run large-scale experiments to drive breakthroughs into production systems.

Required Qualifications

  • Proven experience with developing multimodal models, VLM, and/or vision models
  • Deep, hands-on expertise in model post-training, knowledge distillation (teacher-student training), and/or reinforcement learning (RLHF/RLAIF, PPO/GSPO, and reward modeling)
  • Strong data-centric instincts for building synthetic-data and preference-data pipelines, Model-as-judge generation, data curation and filtering, data augmentations, and/or reasoning about data mixtures and ablations
  • A hands-on builder who enjoys training models, running experiments, debugging pipelines, and integrating ML systems into production while staying grounded in product impact and real-world quality
  • Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow), and the ability to communicate clearly and align research with product and engineering priorities
  • Ability to lead complex research efforts, to communicate clearly and collaborate across teams, while staying grounded in product impact, user experience, and real-world performance

Desired Qualifications

  • Experience with machine translation, multilingual NLP, efficient long-context modeling, language quality estimation, or multimodal machine translation
  • Experience designing evaluation and reward signals using automatic metrics, Model-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation
  • Experience with multi-objective optimization, consistency models, unified multimodal generation
  • Experience with diffusion models
  • Publications at top-tier venues

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce