Senior Performance Engineer
RemoteUnited Kingdom or Germany
Job Summary
Analyze end-to-end performance of large-scale AI workloads across compute, network, storage, and software stacks. Design and execute rigorous performance studies to establish baselines, diagnose regressions, and quantify bottlenecks. Define evaluation methodologies, benchmarks, and success metrics for AI workloads. Use profiling and observability to turn measurements into actionable optimization plans. Partner with deep learning engineers, platform teams, and GPU architects to validate improvements and communicate findings that influence system design decisions. This role supports NVIDIA's DGX Cloud AI Efficiency Team in advancing the performance and resiliency of large-scale AI systems.
Required Qualifications
- BS or higher degree in computer science, computer engineering, or a related field
- 12+ years of experience
- Strong programming skills in C++ and Python
- ability to build reliable analysis and automation workflows
- Solid foundation in operating systems, computer architecture, and distributed systems
- Experience with performance engineering, benchmarking, profiling, and optimization of complex software or systems
- Ability to communicate technical findings, prioritize high-impact work, and build alignment across teams
Desired Qualifications
- Experience analyzing large-scale AI clusters or distributed training and inference workloads
- Experience with CUDA, GPU computing systems, and GPU performance analysis
- Hands-on experience with deep learning frameworks such as PyTorch or JAX/XLA
- Deep understanding of system-level performance analysis, workload characterization, and optimization
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.