Crusoe logo
CrusoePosted 2 weeks ago

Senior Staff Software Engineer, DC Infrastructure

$250,000–$300,000 year

On-siteSan Francisco, California, United States or Sunnyvale, California, United States

Full TimeSenior LevelMediumTechnology

Job Summary

Develop deep-level diagnostics and automation tooling for GPU racks and high-density compute systems, focusing on NVIDIA A100, H200, GB200, B200, and AMD 350X/355X platforms. Build and implement AI agents for component-level diagnosis, remediation, and post-repair validation using burn-in, PyTorch, and NCCL. Own deployment, monitoring, and operational support for fleet management, critical environment tools, and facilities systems including power and direct liquid cooling. Ensure solutions maximize GPU fleet availability and performance to drive customer success.

Required Qualifications

  • Software engineering experience
  • The ability to identify a problem, rapidly develop a scalable solution and ship it
  • Ability to lean in and assist team members working on critical or complex technical initiatives
  • Ability to set the technical direction for a specific project and execute
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)
  • Strength in at least one programming language - Go, Python, Java, Rust
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration skills
  • Ability to work independently and within a team

Desired Qualifications

  • Experience with Temporal and Kubernetes
  • Experience working directly with hardware vendors
  • Background in large-scale GPU fleet operations or hyperscale data center environments

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce