Staff Software Engineer, ML Training and Inference Infrastructure
On-siteLondon, England, United Kingdom
London, England, United KingdomOn-siteFull TimeSenior LevelDoctorate Or Professional DegreeEnterprise
Full TimeSenior LevelDoctorate Or Professional DegreeEnterprise
Job Summary
Optimize Deep Learning training workloads on NVIDIA GPU systems at scale and reduce latency for model inference, pre- and post-processing on onboard systems. Design, train, and deploy large deep learning models leveraging vast labeled and unlabeled datasets. Profile models and perform detective work to improve training and inference speeds, utilizing distributed training frameworks and transformer architecture acceleration techniques.
Required Qualifications
- PhD in CS/CE/EE, or equivalent, in industry experience
- Deep knowledge of PyTorch
- In-depth knowledge of transformer architecture and ways to accelerate the training and inference of transformer models
- Experience of performing large scale distributed training of models
- A track record of profiling models and doing detective work to improve model training and inference speed
Desired Qualifications
- Knowledge of model training framework (e.g. PyTorch Lightning, ray, etc.)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.