Member of Technical Staff — Compute Cluster
On-siteSan Francisco, California, United States
Job Summary
Design, deploy, and operate large distributed GPU clusters end to end, handling provisioning, imaging, upgrades, and capacity planning. Extend scheduling and orchestration systems for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads. Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers. Own cluster storage and artifact paths for checkpoints and logs, while monitoring and continuously improving reliability and error recovery. Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs.
Required Qualifications
- Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
- Strong systems background: Linux, networking, storage, infrastructure-as-code
- Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
- Understanding of monitoring, logging, observability, and version control best practices for ML systems
- Familiarity with CUDA/NCCL and performance profiling for distributed workloads
- Owns deliverables end-to-end, from requirements through autonomous execution
Desired Qualifications
- We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.