NVIDIA logo
NVIDIAPosted 1 month ago

Senior Solutions Architect, AI Infrastructure Enterprise ISVs

$184,000–$287,500 year

RemoteCalifornia, United States or Santa Clara, California, United States

Full TimeSenior LevelDoctorate Or Professional DegreeEnterprise

Job Summary

Partner with ISVs on discovery, architecture reviews, technical deep dives, POCs, benchmarks, and production deployment guidance for AI infrastructure. Advise on the design, build-out, and optimization of accelerated AI clusters across compute, networking, storage, and data center operations. Drive adoption of monitoring and management tools to improve cluster utilization and reliability while building reference architectures, deployment guides, and technical playbooks. Travel up to 20% for customer meetings. This role supports the interdisciplinary team at NVIDIA helping customers adopt accelerated systems for training, fine-tuning, inference, and agentic AI workloads.

Required Qualifications

  • BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
  • 8+ years of hands-on experience in AI infrastructure, accelerated computing, distributed systems, cloud infrastructure, high-performance computing, or machine learning platforms
  • Strong experience designing, deploying, and operating accelerated computing infrastructure at scale
  • In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
  • Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
  • Excellent presentation, communication, problem-solving, documentation, and collaboration skills
  • Travel up to 20%
  • customer meetings may be required

Desired Qualifications

  • Experience architecting AI factories, large GPU clusters, multi-node training environments, production inference platforms
  • Experience deploying LLM training, fine-tuning, RAG, and inference workflows on large-scale AI infrastructure
  • Experience evaluating cluster performance using benchmarks such as MLPerf, HPL, or workload-specific performance tests
  • Applications and systems-level knowledge of OpenMPI, NCCL, distributed training frameworks, and GPU communication patterns
  • Experience delivering technical training, workshops, whitepapers, blogs, or mentoring engineers, researchers, and customers on AI/HPC infrastructure

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce