Staff Hardware Systems Engineer
$215,000–$260,000 year
On-siteSan Francisco, California, United States or Sunnyvale, California, United States
Job Summary
Drive the full hardware development and sustaining lifecycle, including feasibility, bring-up, validation, deployment, and ongoing production support for high-performance compute systems. Develop and maintain scripting and automation frameworks for hardware testing, diagnostics, and continuous reliability improvements. Lead deep troubleshooting and debugging across PCIe, InfiniBand, and NVMe/storage to resolve system-level issues and enable stable production operation. Conduct rigorous system validation and characterization for GPU, CPU, and high-performance compute platforms while collaborating with cross-functional teams to ensure scalable deployment. Support E2E integration and solution testing to meet performance and reliability expectations. Provide data-driven insights to influence the hardware roadmap and reliability strategy.
Required Qualifications
- 8–10+ years of experience in hardware development, validation, sustaining engineering, or production engineering
- Strong hands-on expertise in PCIe, InfiniBand, and NVMe/storage debugging and development
- Deep proficiency in hardware bring-up, board-level debugging, and system-level validation
- Ability to design and implement automation frameworks for hardware testing (Python, Shell, or similar)
- Technical background in digital and analog design, server architecture, and high-performance compute hardware
- Experience working across thermal, mechanical, firmware, and software functions in multidisciplinary environments
- Strong analytical and problem-solving skills with a data-driven approach
- Excellent communication and collaboration skills for working with internal teams and external partners
- Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or equivalent experience
Desired Qualifications
- Experience designing or optimizing GPU-to-GPU communication architectures for AI/ML workloads
- Direct experience integrating NVLink or other next-generation GPU interconnect technologies
- Familiarity with cutting-edge GPU architectures and how to leverage them in AI/HPC environments
- Expertise supporting or designing systems across both ARM and x86 server architectures
- Background in sustainable or energy-efficient hardware design practices
- Advanced certifications or coursework in AI/HPC hardware systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.