Reinforcement Learning Engineer - Ingénieur(e) en apprentissage par renforcement
HybridMontréal, Quebec, Canada
Job Summary
Design robust simulation environments and tune complex reward functions to train autonomous agents within multi-sensor landscapes. Develop and optimize reinforcement learning algorithms such as PPO, SAC, or Offline RL for high-dimensional 3D observation spaces while implementing domain randomization to bridge the reality gap. Collaborate cross-functionally with ML, annotation, and TPM teams to specify data, simulation, and training requirements. This role requires a Master's or PhD in Robotics, Computer Science, or AI, fluency in Python and Git, and experience with frameworks like Ray RLlib or Stable Baselines3. Candidates should have prior industry experience in robotics, game development, or aerospace.
Required Qualifications
- Graduate degree (Master's or PhD) in Robotics, Computer Science, AI, or a related field with a focus on Reinforcement Learning, Imitation Learning, or other Online Machine Learning fields
- Proven experience as an RL Engineer or Research Engineer in a fast-paced environment
- Prior experience in industries with complex multi-disciplinary teams such as robotics, smart grids, precision agriculture, game development, or aerospace
- Fluency with Python, Git, and the Unix shell
- Deep familiarity with frameworks like Ray Rllib, Stable Baselines3, or CleanRL
- Experience with physics engines (MuJoCo, Bullet) or 3D game engines
- Strong Mathematical Background: Essential for understanding Markov Decision Processes (MDPs) and gradient-based optimization
- High Attention to Detail: Critical for debugging non-deterministic agent behaviors and ensuring environment parity
Desired Qualifications
- Experience manipulating virtual environments to train autonomous agents
- Experience designing robust simulation environments, reward structures, and policy architectures that can navigate complex, multi-sensor landscapes
- Experience with tools like Unity, Unreal, or Isaac Sim
- Experience designing and tuning complex reward functions that align agent behavior with product goals and safety constraints
- Experience developing and optimizing RL algorithms (e.g., PPO, SAC, or Offline RL) capable of handling high-dimensional 3D observation spaces
- Experience analyzing the 'reality gap' and implementing domain randomization or adaptation techniques to ensure models perform reliably in real-world scenarios
- Familiarity with collaborative tools such as Jira/Confluence, Slack, a Git server, and an experiment tracking framework
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.