Sarvam AI logo
Sarvam AIPosted 2 months ago

Infrastructure SRE - HPC

On-siteBengaluru, Karnataka, India or Chennai, Tamil Nadu, India

Full TimeAssociates DegreeSmall

Job Summary

Infrastructure SRE - HPC at Sarvam overseeing a large, multi-vendor GPU fleet supporting training jobs spanning hundreds of GPUs and inference services with strict latency guarantees. The role emphasizes reliability challenges above Kubernetes administration, including parallel filesystems under heavy checkpoint load, RDMA fabric issues, NCCL hangs, driver/firmware drift, and distributed training failures. The team seeks specialists with deep expertise in one focus area and working fluency across the others, ready to manage end-to-end GPU fleet operations, on-call rotations, runbooks, and durable postmortems. You will operate the GPU fleet end to end (provisioning, observability, capacity, fleet health), own on-call rotations, build internal tooling, and collaborate with ML and platform teams to keep large runs alive and latency predictable. Required qualifications include 5+ years in infrastructure or SRE, 2+ years operating GPU clusters at scale, proficiency in Python or Go, on-call ownership, and the ability to triage and route problems across the five focus areas; bonus points for Slurm/Kubernetes hybrid environments, on-prem GPU deployment, and experience with Indian NCPs, DGX SuperPOD, Lambda, and related stacks.

Required Qualifications

  • 5+ years in infrastructure or site reliability engineering
  • 2+ years operating GPU clusters at scale
  • Proficiency in Python or Go
  • On-call ownership of infrastructure
  • Working fluency across multiple focus areas (Storage, Fabric, etc.)

Desired Qualifications

  • Slurm and Kubernetes hybrid environments.
  • On-premise GPU deployment, including coordination with datacenter operations on power, cooling, and InfiniBand cabling.
  • Experience with Indian NCPs, DGX SuperPOD, Lambda, CoreWeave, NeevCloud etc.
  • Multi-tenant GPU isolation (MIG, MPS, time-slicing) in production.
  • Why Sarvam?
  • Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.
  • Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar
  • High ownership and high impact, from day one
  • Everything we do is AI-first, from the way we build and ship to the way we think about problems
  • You can work on problems that could change how an entire country learns, works, and communicates
  • If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce