Senior Machine Learning Engineer, AI Performance
On-siteLondon, England, United Kingdom
Job Summary
Own end-to-end delivery of model releases from initial requirements through training, evaluation, iteration, and final readiness for deployment. Train and iterate on PyTorch models with a hypothesis-driven approach, debugging performance regressions and root-causing issues to maintain tight runtime constraints. Apply optimization techniques like quantization and distillation to balance model capability with latency, memory, and power requirements. Collaborate cross-functionally with performance engineering teams to define bottlenecks and align on optimization priorities. Communicate clearly with stakeholders regarding delivery timelines and readiness criteria.
Required Qualifications
- Proven experience improving performance in production systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost)
- Strong hands-on experience training and iterating on deep learning models in PyTorch (not just using high-level tooling)
- Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly
- Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime execution
- Familiarity with model optimisation concepts such as quantisation and/or distillation (hands-on is a strong signal, but not a strict requirement if the fundamentals are solid)
- Ability to reason across multiple levels of abstraction—from high-level model behaviour down to practical runtime/latency implications
- Strong engineering fundamentals and collaboration skills
Desired Qualifications
- Experience working on models that must meet tight latency / efficiency constraints (edge, embedded, real-time, or similarly constrained production settings)
- Exposure to ML systems spanning training → evaluation → deployment handoff (even if you're not writing kernels day-to-day)
- Exposure to embedded or edge deployment of ML models, including benchmarking on real devices and handling system-level constraints
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.