Machine Learning Performance Engineer
On-siteLondon, England, United Kingdom
Job Summary
Design and implement techniques to improve performance and capabilities of research workloads on distributed CPU, GPU, and memory-intensive jobs. Collaborate with researchers, senior stakeholders, and engineers to profile, benchmark, and tune large-scale training and inference workloads while eliminating bottlenecks. Develop reference implementations, libraries, and tools to enhance job efficiency and reliability, directly influencing long-term platform evolution and architecture decisions. Work closely with systems, architecture, and platform teams to evolve the compute stack and support innovation across the firm.
Required Qualifications
- Bachelors, Masters or PhD degree in computer science, or equivalent experience
- Proven track record of profiling, benchmarking and optimising distributed workloads
- Experience with Python
- Knowledge of CUDA
- Experience with HPC schedulers and Kubernetes-based workload orchestration
- Strong understanding of one or more deep learning frameworks, such as PyTorch
- Strong background in data structures, algorithms, and parallel programming on heterogeneous systems
- Deep understanding of Linux OS fundamentals, such as as scheduling, memory management, NUMA, networking, and filesystems
- Familiarity with profiling and monitoring tools, such as nsys, ncu, eBPF-based tools, and performance counters
- Strong communication skills with the ability to collaborate across research, infrastructure and engineering teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.