Research Scientist / Engineer – Reinforcement Learning Infrastructure
$187,500–$395,000 year
RemoteUnited Kingdom or United States
Job Summary
Design, build, and scale distributed RL post-training systems for large multimodal models by orchestrating trainer, rollout, environment, and reward workloads across thousands of GPUs. Build and optimize high-throughput rollout generation, integrating inference engines like vLLM and SGLang into the training loop with efficient weight synchronization and asynchronous schemes. Design and implement RL environments for agentic and multi-step tasks, including sandboxed code execution and multimodal interaction, ensuring reproducibility and scalability to millions of episodes. Develop reward infrastructure featuring verifiable rewards, reward model serving, and LLM-as-judge pipelines while maintaining defenses against reward hacking. Advance training efficiency through sequence packing, KV cache reuse, and resource scheduling across heterogeneous workloads. Collaborate with researchers to translate new post-training ideas into production-quality training runs.
Required Qualifications
- Hands-on experience post-training LLMs with reinforcement learning (e.g. PPO / GRPO-family methods, RLHF, RLVR / RL from verifiable rewards) at meaningful scale
- Extensive experience with distributed PyTorch training and parallelization strategies (FSDP, Tensor / Pipeline / Expert Parallel) for foundation models
- Experience building RL environments, reward functions, verifiers, or evaluation harnesses for LLM agents — including sandboxed execution and multi-turn tool use
- Deep familiarity with RL post-training frameworks and their systems tradeoffs (e.g. veRL, OpenRLHF, TRL, Ray-based orchestration) and inference engines used for rollouts (vLLM, SGLang)
- Strong understanding of GPU clusters, networking, and communication libraries (NCCL, MPI), and how they behave under mixed training + inference workloads
Desired Qualifications
- (Preferred) Experience running RL training across >100 GPUs, including asynchronous or disaggregated trainer/rollout architectures
- (Preferred) Experience with containerization and orchestration (Kubernetes, Ray) for large environment fleets and sandboxed workloads
- (Preferred) Research contributions in RL for LLMs — reasoning, agents, reward modeling, or long-horizon tasks — or open-source contributions to RL training frameworks
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.