Crusoe logo
CrusoePosted 1 month ago

Principal Engineer, CAPE

$285,000–$335,000 year

On-siteSan Francisco, California, United States

Full TimeSenior LevelMediumTechnology

Job Summary

Build unified observability planes correlating GPU, networking, storage, and workload signals to enable fast diagnosis and recovery. Own the fleet as a single logical computer by designing health models, schedulers, and source-of-truth systems that treat tens of thousands of accelerators as one programmable unit. Implement closed-loop autonomy to diagnose, decide, and remediate failures without human intervention, while forecasting hardware degradation hours in advance to preemptively migrate workloads. Maximize goodput as the primary objective function by optimizing scheduling, placement, and maintenance decisions against real-time energy availability and thermal headroom. Design digital twins to simulate failures and remediation logic before production deployment, and develop agentic operations that propose and execute fixes with full audit trails. Establish self-qualifying hardware pipelines where new and repaired nodes prove themselves through automated burn-in.

Required Qualifications

  • 10+ years building infrastructure-layer systems at scale
  • Deep experience with distributed systems design
  • Hands-on fluency with GPU/HPC infrastructure
  • Track record of designing and shipping large-scale observability or telemetry platforms
  • Comfort operating in ambiguity and defining the architecture and standards for a system that doesn't exist yet
  • Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar)

Desired Qualifications

  • Experience applying ML/statistical methods to noisy operational telemetry
  • Prior exposure to zero-trust or policy-based multi-tenancy architectures

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce