NVIDIA logo
NVIDIAPosted 1 month ago

Senior Solutions Architect, CSP System

On-siteShanghai, Shanghai, China

Full TimeSenior LevelEnterprise

Job Summary

Partner with Sales, BD, and CPM teams to land NVIDIA GPU and AI Infra technologies into top-tier Chinese CSP accounts. Serve as the primary technical authority for end-to-end consultation on GPU cluster architecture, AI workload deployment, and full-stack software optimization. Conduct in-depth bottleneck analysis and implement system-level tuning for AI training, inference, RL, and gaming workloads. Lead open-source contributions for NVIDIA AI infra stacks and shape industry technical standards. Act as the key liaison between Chinese CSP customers and global engineering teams to drive product roadmap iteration and ensure export compliance. Quantify business value through technical workshops, PoCs, and production pilots to accelerate large-scale replication. Mentor junior SA team members and standardize solution delivery processes.

Required Qualifications

  • Bachelor's/Master's/PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field
  • 8+ years of hands-on experience in GPU architecture, AI system optimization, large-scale data center infrastructure, or hyperscale cloud computing
  • Solid experience in AI training/inference, distributed computing or HPC workloads
  • Deep understanding of GPU microarchitecture, CUDA programming model, GPU memory hierarchy and system scheduling mechanisms
  • Proficient in performance profiling, bottleneck analysis and end-to-end AI workload tuning
  • Strong programming proficiency in C/C++ and Python
  • Familiar with CUDA kernels, compiler toolchains, AI framework optimization (PyTorch/TensorRT) and large-scale distributed system tuning
  • Proven hands-on experience working with major Chinese CSPs or global hyperscalers
  • In-depth knowledge of their public cloud AI service architectures, cluster operation mechanisms and core workload characteristics
  • Excellent technical communication and presentation skills
  • Capable of explaining complex GPU system and AI infra technologies to technical engineers, architecture teams and business stakeholders
  • Strong cross-functional collaboration capability
  • Able to work efficiently in a global matrix team
  • Prioritize multiple high-value technical projects under fast-paced business demands
  • Hands-on engineering capability
  • Result-oriented, self-driven
  • Able to independently own end-to-end technical project delivery
  • Committed, proactive, and capable of sustaining high-quality technical output for long-term strategic CSP projects
  • Familiar with NVIDIA full-stack products (GPU data center hardware, TensorRT-LLM, Dynamo, NCCL, CUDA software stack)
  • Hands-on experience with Vera/Grace CPU + GPU heterogeneous co-optimization
  • Familiar with AI agent and RL training system tuning
  • In-depth experience in Dynamo LLM inference optimization, including KV Cache management, intelligent scheduling planner and dynamic resource scaling
  • Open-source contribution experience in AI infra, GPU optimization libraries, or distributed computing frameworks with public upstream records
  • Solid experience in Agentic AI, RL post-training or long-context LLM workload optimization on GPU clusters
  • Familiar with semiconductor and data center technology export compliance requirements in China market
  • Proven track record of independently leading CSP technical PoC, pilot verification and large-scale production deployment projects with measurable business outcomes

Desired Qualifications

  • equivalent industry experience is highly valued
  • equivalent industry experience is highly valued

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce