Senior Software Engineer - Managed Kubernetes
$266,000–$395,000 year
HybridSan Francisco, California, United States or San Jose, California, United States
Job Summary
Design and maintain scalable control plane services, operators, and custom Kubernetes controllers for AI workloads, developing automation in Go/Python for cluster lifecycle management. Build GPU-aware orchestration systems supporting scheduling and resource allocation, and partner with the Network team on CNI integration, high-performance fabrics, and RDMA. Write resilient systems handling failure gracefully across distributed environments and develop platform services for inference, including model serving and autoscaling. Support and debug production issues through on-call rotation. Requires 6+ years of software engineering experience with deep Kubernetes internals knowledge, distributed systems fundamentals, and strong Go/Python skills. Full-time presence in San Francisco, San Jose, or Bellevue office four days per week.
Required Qualifications
- 6+ years of experience in software engineering, with a track record of owning significant technical scope within a team (e.g., driving a project from design through production, or acting as a de facto tech lead on a workstream)
- Deep understanding of Kubernetes internals: controllers, schedulers, operators, CRDs, CSI, CNI, and the extension patterns that make Kubernetes powerful
- Solid grasp of distributed systems fundamentals — fault tolerance, graceful degradation, and failure handling in large-scale environments
- Experience operating the control plane and low-level pieces of large-scale Kubernetes clusters
- Experience with observability at scale: Prometheus, Grafana, distributed tracing, and building actionable alerting systems
- Strong programming skills in Go and Python; ability to collaborate effectively on shared codebases
- Solid knowledge of Linux systems, networking, containers, and cloud infrastructure
- Take pride in owning and delivering core components of products and platforms
- Must be able to work in our San Francisco, San Jose, or Bellevue office location 4 days per week
Desired Qualifications
- Experience building and operating managed Kubernetes services (GKE, EKS, AKS, or similar) or working on Kubernetes control plane components
- Hands-on experience with NVIDIA's GPU/networking ecosystem: GPU Operator, device plugins, DCGM, MIG, Network Operator, NCCL tuning, or similar
- Familiarity with HPC and traditional job schedulers (Slurm) and Kubernetes-native batch scheduling (KAI, Volcano, Kueue)
- Familiarity with GPU, InfiniBand, RDMA, or high-performance computing on Kubernetes
- Exposure to storage architecture for AI/ML workloads
- Past contributions to CNCF projects or Kubernetes SIGs a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.