Senior AI Training Performance Engineer
On-siteShanghai, Shanghai, China
Job Summary
Analyze, profile, and optimize AI and deep learning training workloads on state-of-the-art hardware and software platforms. Prioritize and solve performance problems across dozens of neural networks by implementing production-quality software in NVIDIA's deep learning platform stack, from drivers to DL frameworks. Build tools to automate workload analysis and optimization while implementing key training workloads in proprietary simulators for future architecture studies. Requires a PhD or MS with extensive experience in GPU architecture, C++, CUDA, and deep learning training. Join the Deep Learning Architecture team to directly impact the hardware and software roadmap in NVIDIA's AI revolution.
Required Qualifications
- PhD (or equivalent experience) in CS, EE or CSEE
- 5+ years
- MS and 8+ years of relevant work experience
- Strong background in deep learning and neural networks, in particular training
- Deep understanding of computer architecture
- Familiarity with the fundamentals of GPU architecture
- Proven experience analyzing and tuning application performance
- Experience with processor and system-level performance modelling
- Programming skills in C++, Python, and CUDA
- Fluency in English
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.