Staff Software Engineer (Cloud Infrastructure)
$215,000–$260,000 year
HybridSan Francisco, California, United States
Job Summary
Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems. Develop software to troubleshoot NVIDIA A100, H200, GB200, B200, and AMD 350X/355X platforms, executing component-level remediation for failed or degraded hardware. Partner with data center operations to manage field-replaceable unit repairs for GPUs, power supplies, cooling systems, and networking hardware. Conduct post-repair validation, burn-in testing, and firmware upgrades while maintaining detailed documentation in ticketing systems. Collaborate with engineering teams to identify root causes of systemic failures and implement preventative solutions. Participate in a rotating infrastructure on-call schedule with daytime coverage and handoff to the Europe team.
Required Qualifications
- Ability to code in Golang
- Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments
- Deep understanding of GPU architectures and hands-on experience with GPU-based systems
- Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms
- Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE)
- Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing
- Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities
- Experience working with enterprise server hardware, power delivery, and cooling systems
- Strong analytical and problem-solving skills
- Excellent communication and collaboration skills
- Ability to work independently in a fast-paced data center or operations environment
- Ability to code in Golang
- Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments
- Deep understanding of GPU architectures and hands-on experience with GPU-based systems
- Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms
- Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE)
- Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing
- Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities
- Experience working with enterprise server hardware, power delivery, and cooling systems
- Strong analytical and problem-solving skills
- Excellent communication and collaboration skills
- Ability to work independently in a fast-paced data center or operations environment
Desired Qualifications
- Technical certification or Associate's/Bachelor's degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience
- Experience working directly with hardware vendors and escalations
- Background in large-scale GPU fleet operations or hyperscale data center environments
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.