Senior Solutions Architect, AI Infrastructure Enterprise ISVs
$184,000–$287,500 year
RemoteCalifornia, United States or Santa Clara, California, United States
Job Summary
Partner with ISVs on discovery, architecture reviews, technical deep dives, POCs, benchmarks, and production deployment guidance for AI infrastructure. Advise on the design, build-out, and optimization of accelerated AI clusters across compute, networking, storage, and data center operations. Drive adoption of monitoring and management tools to improve cluster utilization and reliability while building reference architectures, deployment guides, and technical playbooks. Travel up to 20% for customer meetings. This role supports the interdisciplinary team at NVIDIA helping customers adopt accelerated systems for training, fine-tuning, inference, and agentic AI workloads.
Required Qualifications
- BS, MS, or PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, other Engineering or related fields (or equivalent experience)
- 8+ years of hands-on experience in AI infrastructure, accelerated computing, distributed systems, cloud infrastructure, high-performance computing, or machine learning platforms
- Strong experience designing, deploying, and operating accelerated computing infrastructure at scale
- In-depth knowledge of AI cluster orchestration, scheduling, automation and CI/CD deployment pipelines
- Understanding of data center networking technologies such as InfiniBand, Ethernet, RDMA, network configuration or performance tuning
- Excellent presentation, communication, problem-solving, documentation, and collaboration skills
- Travel up to 20%
- customer meetings may be required
Desired Qualifications
- Experience architecting AI factories, large GPU clusters, multi-node training environments, production inference platforms
- Experience deploying LLM training, fine-tuning, RAG, and inference workflows on large-scale AI infrastructure
- Experience evaluating cluster performance using benchmarks such as MLPerf, HPL, or workload-specific performance tests
- Applications and systems-level knowledge of OpenMPI, NCCL, distributed training frameworks, and GPU communication patterns
- Experience delivering technical training, workshops, whitepapers, blogs, or mentoring engineers, researchers, and customers on AI/HPC infrastructure
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.