NVIDIA logo
NVIDIAPosted 1 month ago

Senior Systems Software Engineer, Accelerated Kubernetes Performance and Scale - DGX Cloud

$152,000–$241,500 year

RemoteUnited States

Full TimeSenior LevelEnterprise

Job Summary

Lead end-to-end performance and scalability analysis across the Kubernetes-based accelerated runtime stack, including NVIDIA components like GPU Operator and node-feature-discovery, tracking issues from orchestration down to the metal. Design upstream architectural changes to the Kubernetes control plane to enable reliable operation at hyperscale cluster sizes and improve container startup latency for smooth inference scaling. Assess and contribute to open-source projects such as Grove and gateway-api-inference-extension, while advancing scalability for confidential containers on Kubernetes. Use DSX simulation infrastructure to model full AI-factory deployments and validate performance across thousands of simulated GPUs, integrating continuous testing into CI/CD workflows. Collaborate with researchers and upstream communities to design automated workload tests and present findings at industry events like KubeCon and GTC.

Required Qualifications

  • Bachelor's or Master's degree in Engineering or equivalent experience, ideally in Electrical, Computer Engineering, or Computer Science
  • 5+ years of experience in computer architecture, networking, storage systems, and accelerator-based platforms
  • Expertise in Kubernetes and familiarity with the broader CNCF ecosystem
  • Deep experience with large-scale, parallel, distributed accelerator systems and performance optimization of AI workloads
  • Experience with performance modeling and benchmarking for large-scale systems
  • Proficiency in Golang and/or Python
  • Strong familiarity with the NVIDIA software stack across training and inference
  • Expertise with at least one major public cloud provider (for example, AWS, Azure, GCP, or OCI)

Desired Qualifications

  • PhD or equivalent experience in relevant areas
  • Strong operational experience with any one of the Kubernetes distributions
  • Prior experience scaling Kubernetes clusters to ultra-large node and object counts
  • Demonstrated history of working in the open-source community
  • Excellent communication and interpersonal abilities

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce