Lambda logo
LambdaPosted 1 month ago

Senior HPC Systems Validation Engineer

$255,000–$340,000 year

HybridSan Jose, California, United States

Full TimeSenior LevelSmall

Job Summary

Own system integration validation for new HPC AI/ML, general purpose compute, storage, and network hardware platforms throughout hardware NPI and deployment readiness. Execute functional, stress, reliability, error-injection, and benchmark tests on new hardware systems, enabling automated tools for scale testing. Collaborate with deployment, fleet engineering, and PMO teams to facilitate tooling automation and L10/L11/L12 level benchmarking. Manage firmware and software compatibility, track releases during NPI, and support maintenance after production. Support RMA team on critical hardware triage and root cause analysis, driving failure correlation and trend analysis to improve product reliability.

Required Qualifications

  • 5 years of experience in one or many of the following areas: hardware integration validation, fleet hardware reliability engineering, component qualification, performance validation for HPC, data center, or cloud infrastructure products
  • Possess deep knowledge in system integration testing and performance benchmarking at one or many of the L10, L11 and L12 levels
  • Experiences in validation of one or many of the following server hardware platforms: AI/ML, general compute (x86 and ARM), storage systems, or network switches
  • Hands-on experiences with vendor-led product NPI cycles and can drive hardware validation, issues triage, debugging and root cause in both NPI and after production at scale
  • Collaborate well across architecture, hardware and datacenter engineering, supply chain, fleet and infrastructure engineering, HPC deployments and operations, and vendor engineering teams to deliver complete, production-ready hardware solutions
  • Strong ownership and can do attitude; self-starter who feels comfortable working in ambiguity
  • Presence in our San Jose office location 4 days per week

Desired Qualifications

  • 10+ years of experience in one or many of the following areas: hardware integration validation, fleet hardware reliability engineering, component qualification, performance validation for HPC, data center, or cloud infrastructure products
  • Experience supporting AI/ML infrastructure and accelerated compute hardware (e.g., NVIDIA, AMD, Intel)
  • Experience in acceptance testing to enable large scale hardware deployments
  • Deep knowledge in BMC, BIOS settings and network card configurations
  • Proficient in validating, issue triaging, benchmarking and performance tuning for rack-scale servers at L10, L11, L12 levels
  • Working knowledge of infrastructure toolings for asset management, provisioning, hardware lifecycle management, etc.
  • Working knowledge in PLM systems and BOM structure

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce