Senior Staff Software Engineer, DC Infrastructure
$250,000–$300,000 year
On-siteSan Francisco, California, United States or Sunnyvale, California, United States
Job Summary
Develop deep-level diagnostics and automation tooling for GPU racks and high-density compute systems, focusing on NVIDIA A100, H200, GB200, B200, and AMD 350X/355X platforms. Build and implement AI agents for component-level diagnosis, remediation, and post-repair validation using burn-in, PyTorch, and NCCL. Own deployment, monitoring, and operational support for fleet management, critical environment tools, and facilities systems including power and direct liquid cooling. Ensure solutions maximize GPU fleet availability and performance to drive customer success.
Required Qualifications
- Software engineering experience
- The ability to identify a problem, rapidly develop a scalable solution and ship it
- Ability to lean in and assist team members working on critical or complex technical initiatives
- Ability to set the technical direction for a specific project and execute
- Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)
- Strength in at least one programming language - Go, Python, Java, Rust
- Strong analytical and problem-solving skills
- Excellent communication and collaboration skills
- Ability to work independently and within a team
Desired Qualifications
- Experience with Temporal and Kubernetes
- Experience working directly with hardware vendors
- Background in large-scale GPU fleet operations or hyperscale data center environments
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.