Causal Labs logo
Causal LabsPosted 1 month ago

Member of Technical Staff — Compute Cluster

On-siteSan Francisco, California, United States

Full TimeSenior LevelStartup

Job Summary

Design, deploy, and operate large distributed GPU clusters end to end, handling provisioning, imaging, upgrades, and capacity planning. Extend scheduling and orchestration systems for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads. Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers. Own cluster storage and artifact paths for checkpoints and logs, while monitoring and continuously improving reliability and error recovery. Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs.

Required Qualifications

  • Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
  • Strong systems background: Linux, networking, storage, infrastructure-as-code
  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
  • Understanding of monitoring, logging, observability, and version control best practices for ML systems
  • Familiarity with CUDA/NCCL and performance profiling for distributed workloads
  • Owns deliverables end-to-end, from requirements through autonomous execution

Desired Qualifications

  • We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce