Hewlett Packard Enterprise logo
Hewlett Packard EnterprisePosted 3 weeks ago

High Performance Compute Systems Site Lead (Onsite - LANL)

On-siteAll, Catalunya, Kingdom of Spain

Full TimeSenior LevelEnterprise

Job Summary

Lead onsite HPC and AI system operations at Los Alamos National Laboratory, providing day-to-day technical leadership for HPE hardware engineers, Linux administrators, and software analysts supporting large-scale Cray environments. Establish daily priorities, coordinate maintenance windows, and manage major incidents while ensuring operational readiness and service quality for mission-critical computing infrastructure. Serve as the primary technical focal point for escalations, root-cause analysis, and customer governance, partnering with the DSM on executive-level issues and service-delivery risks. Perform hands-on diagnostics using Linux command-line tools, hardware telemetry, and out-of-band management interfaces to troubleshoot compute nodes, storage, and network components, including rack-level hardware replacement and cable management. Mentor team members, facilitate operational reviews, and maintain accurate documentation and site procedures in a 24x7 production environment requiring onsite presence Monday through Friday.

Required Qualifications

  • US Citizenship and the ability to obtain and maintain a DOE Q Clearance
  • Must work onsite M-F in Los Alamos, New Mexico, with additional onsite work as required for planned maintenance, major incidents, and on-call support
  • This is not a remote or hybrid position
  • High school diploma or equivalent with at least 7 years of relevant technical experience
  • Associate or bachelor's degree in a technical field with at least 5 years of relevant technical experience
  • 5+ years of hands-on experience supporting complex electronic systems, enterprise server hardware, integrated data center infrastructure, or comparable production technology environments
  • Experience must include diagnosing, repairing, or maintaining enterprise server components, including processors, memory, storage devices, power supplies, BMCs, network adapters, optical connectivity, and copper or fiber cabling
  • 3+ years of hands-on experience supporting HPC systems, supercomputing environments, large-scale Linux clusters, or similarly complex Linux-based compute infrastructure
  • Experience must include troubleshooting interactions among compute nodes, management systems, high-speed interconnects, storage platforms, operating systems, workload managers or schedulers, power, cooling, and supporting infrastructure
  • 3+ years of experience providing technical leadership, mentoring, work coordination, or task direction for a multidisciplinary technical team
  • Direct people-management experience is not required
  • 3+ years of hands-on Linux system administration, production support, or troubleshooting experience with Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or a comparable enterprise Linux distribution
  • Candidates must be able to independently use Linux command-line tools to navigate files

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce