G-Research logo
G-ResearchPosted 1 week ago

Machine Learning Performance Engineer

On-siteLondon, England, United Kingdom

Full TimeMedium

Job Summary

Design and implement techniques to improve performance and capabilities of research workloads on distributed CPU, GPU, and memory-intensive jobs. Collaborate with researchers, senior stakeholders, and engineers to profile, benchmark, and tune large-scale training and inference workloads while eliminating bottlenecks. Develop reference implementations, libraries, and tools to enhance job efficiency and reliability, directly influencing long-term platform evolution and architecture decisions. Work closely with systems, architecture, and platform teams to evolve the compute stack and support innovation across the firm.

Required Qualifications

  • Bachelors, Masters or PhD degree in computer science, or equivalent experience
  • Proven track record of profiling, benchmarking and optimising distributed workloads
  • Experience with Python
  • Knowledge of CUDA
  • Experience with HPC schedulers and Kubernetes-based workload orchestration
  • Strong understanding of one or more deep learning frameworks, such as PyTorch
  • Strong background in data structures, algorithms, and parallel programming on heterogeneous systems
  • Deep understanding of Linux OS fundamentals, such as as scheduling, memory management, NUMA, networking, and filesystems
  • Familiarity with profiling and monitoring tools, such as nsys, ncu, eBPF-based tools, and performance counters
  • Strong communication skills with the ability to collaborate across research, infrastructure and engineering teams

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce