Senior Research Scientist | Model Steering
HybridLondon, England, United Kingdom
Job Summary
Lead fine-tuning, post-training, and reinforcement learning for DeepL's next-generation LLM-based translation models, focusing on making systems steerable to user preferences and rules. Fuse high-value human expert data with synthetic data to build reward and evaluator models for translation, while driving hands-on R&D on supervised fine-tuning, knowledge distillation, and preference optimization. Own the full lifecycle of model delivery from prototyping and ablation to production deployment at scale, establishing practices for evaluation, reproducibility, and continuous improvement. Mentor researchers and engineers to raise model quality standards within a global team of over 1,000 employees.
Required Qualifications
- Proven experience making large models steerable and instruction-following by identifying the most effective method to instill a given behavior, drawing from instruction tuning, latent space methods, steering vectors, and/or constrained encoding and decoding methods.
- Deep, hands-on expertise in LLM post-training (SFT, DPO), knowledge distillation (teacher-student training), and/or reinforcement learning (RLHF/RLAIF, PPO/GSPO, and reward modeling).
- Strong data-centric instincts for building synthetic-data and preference-data pipelines, LLM-as-judge generation, data curation and filtering, and reasoning about data mixtures and ablations.
- Experience designing evaluation and reward signals using automatic metrics, LLM-as-judge evaluation, non-verifiable rewards, and human-in-the-loop evaluation.
- A hands-on builder who enjoys training models, running experiments, debugging pipelines, and integrating ML systems into production while staying grounded in product impact and real-world quality.
- Ownership of a substantial research direction with strong execution, and experience mentoring others on a fast-moving, applied research team.
- Strong coding and experimentation skills (Python, PyTorch/JAX/Tensorflow), and the ability to communicate clearly and align research with product and engineering priorities.
Desired Qualifications
- Demonstrated experience fine-tuning and training large models at scale, including distributed/multi-node training (e.g. FSDP, DeepSpeed, or Megatron-style frameworks) and efficient training techniques.
- Experience fine-tuning existing reasoning models for specific tasks and behaviors without degrading their reasoning capabilities.
- Experience with machine translation, multilingual NLP, or language quality estimation.
- Familiarity with inference and serving at scale (e.g. via vLLM, SGLang, TensorRT-LLM, etc) and long-context modelling.
- Publications at top-tier venues.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.