Research Engineer/Scientist - Machine Learning RL & Optimisation (Contractor)
On-siteLondon, England, United Kingdom
Job Summary
Design and execute scaled RL finetuning workflows (e.g., PPO, GRPO) to enhance LLM reasoning, instruction-following, and alignment. Architect and manage large-scale distributed training experiments across multi-node GPU, optimizing for maximum throughput and hardware utilization. Develop and maintain training infrastructure using advanced parallelization frameworks (verl, trl, DeepSpeed, FSDP) to support rapidly evolving research needs. Integrate high-performance inference engines like vLLM directly into RL generation loops to reduce rollout latency and accelerate training cycles. Implement robust profiling and debugging pipelines to diagnose bottlenecks in GPU memory, compute, and inter-node communication. Collaborate with data and evaluation teams to design dense reward functions and synthetic data generation pipelines. Design, benchmark, and deploy highly optimized custom tensor operators (e.g., FlashAttention, GEMM) across heterogeneous hardware architectures using modern Domain-Specific Languages (DSLs) and AI compilers.
Required Qualifications
- Master's or PhD (or equivalent industry research experience) in Machine Learning, Computer Science, Data Science, or a highly quantitative field with a heavy focus on Machine Learning
- Deep proficiency in PyTorch and experience writing custom training loops and data pipelines
- Hands-on experience with RLHF/RLVF methods (e.g., PPO, GRPO) and an understanding of policy optimization dynamics
- Technical familiarity with at least two of the following: DeepSpeed, FSDP, verl or trl
- Production or heavy research experience utilizing vLLM or similar high-throughput inference serving engines for generation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.