Lambda logo
LambdaPosted 1 week ago

Senior Software Engineer - Managed Kubernetes

$266,000–$395,000 year

HybridSan Francisco, California, United States or San Jose, California, United States

Full TimeSenior LevelSmall

Job Summary

Design and maintain scalable control plane services, operators, and custom Kubernetes controllers for AI workloads, developing automation in Go/Python for cluster lifecycle management. Build GPU-aware orchestration systems supporting scheduling and resource allocation, and partner with the Network team on CNI integration, high-performance fabrics, and RDMA. Write resilient systems handling failure gracefully across distributed environments and develop platform services for inference, including model serving and autoscaling. Support and debug production issues through on-call rotation. Requires 6+ years of software engineering experience with deep Kubernetes internals knowledge, distributed systems fundamentals, and strong Go/Python skills. Full-time presence in San Francisco, San Jose, or Bellevue office four days per week.

Required Qualifications

  • 6+ years of experience in software engineering, with a track record of owning significant technical scope within a team (e.g., driving a project from design through production, or acting as a de facto tech lead on a workstream)
  • Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
  • Solid grasp of distributed systems fundamentals — fault tolerance, graceful degradation, and failure handling in large-scale environments
  • Experience operating the control plane and low-level pieces of large-scale Kubernetes clusters
  • Experience with observability at scale: Prometheus, Grafana, distributed tracing, and building actionable alerting systems
  • Strong programming skills in Go and Python; ability to collaborate effectively on shared codebases
  • Solid knowledge of Linux systems, networking, containers, and cloud infrastructure
  • Take pride in owning and delivering core components of products and platforms
  • Must be able to work in our San Francisco, San Jose, or Bellevue office location 4 days per week

Desired Qualifications

  • Experience building and operating managed Kubernetes services (GKE, EKS, AKS, or similar) or working on Kubernetes control plane components
  • Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning, or similar
  • Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
  • Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes
  • Exposure to storage architecture for AI/ML workloads
  • Past contributions to CNCF projects or Kubernetes SIGs a plus

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce