Rivian logo
RivianPosted 14 months ago

Staff Software Engineer, ML Training and Inference Infrastructure

On-siteLondon, England, United Kingdom

Full TimeSenior LevelDoctorate Or Professional DegreeEnterprise

Job Summary

Optimize Deep Learning training workloads on NVIDIA GPU systems at scale and reduce latency for model inference, pre- and post-processing on onboard systems. Design, train, and deploy large deep learning models leveraging vast labeled and unlabeled datasets. Profile models and perform detective work to improve training and inference speeds, utilizing distributed training frameworks and transformer architecture acceleration techniques.

Required Qualifications

  • PhD in CS/CE/EE, or equivalent, in industry experience
  • Deep knowledge of PyTorch
  • In-depth knowledge of transformer architecture and ways to accelerate the training and inference of transformer models
  • Experience of performing large scale distributed training of models
  • A track record of profiling models and doing detective work to improve model training and inference speed

Desired Qualifications

  • Knowledge of model training framework (e.g. PyTorch Lightning, ray, etc.)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce