NVIDIA logo
NVIDIAPosted 3 weeks ago

Senior Platform Engineer, Network Infrastructure

RemoteBengaluru, Karnataka, India or Pune, Maharashtra, India

Full TimeSenior LevelBachelors DegreeEnterprise

Job Summary

Design, build, and operate the Kubernetes platform powering NVIDIA's global network automation, telemetry, and operations across data centers, colocation, and cloud environments. Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, and recovery. Develop production-quality software and automation for cluster provisioning, validation, remediation, and safe multi-cluster delivery through GitOps. Provide production support for network services hosted on the platform, diagnosing complex failures involving control-plane health, networking, storage, and scheduling. Drive issues from initial signal through verified resolution and lead incident response and recovery. Participate in CFR's production on-call rotation, including scheduled after-hours and weekend coverage, while establishing consistent engineering practices across US and Bangalore teams.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience
  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems
  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery
  • Proficiency in at least one general-purpose programming language, such as Go or Python
  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery
  • Experience deploying and supporting network automation or telemetry services on Kubernetes
  • Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion

Desired Qualifications

  • Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus
  • Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery
  • Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades
  • Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns
  • Experience designing or operating network automation and telemetry services on Kubernetes at global scale
  • Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce