High Performance Compute Systems Site Lead (Onsite - LANL)
On-siteAll, Catalunya, Kingdom of Spain
Job Summary
Lead onsite HPC and AI system operations at Los Alamos National Laboratory, providing day-to-day technical leadership for HPE hardware engineers, Linux administrators, and software analysts supporting large-scale Cray environments. Establish daily priorities, coordinate maintenance windows, and manage major incidents while ensuring operational readiness and service quality for mission-critical computing infrastructure. Serve as the primary technical focal point for escalations, root-cause analysis, and customer governance, partnering with the DSM on executive-level issues and service-delivery risks. Perform hands-on diagnostics using Linux command-line tools, hardware telemetry, and out-of-band management interfaces to troubleshoot compute nodes, storage, and network components, including rack-level hardware replacement and cable management. Mentor team members, facilitate operational reviews, and maintain accurate documentation and site procedures in a 24x7 production environment requiring onsite presence Monday through Friday.
Required Qualifications
- US Citizenship and the ability to obtain and maintain a DOE Q Clearance
- Must work onsite M-F in Los Alamos, New Mexico, with additional onsite work as required for planned maintenance, major incidents, and on-call support
- This is not a remote or hybrid position
- High school diploma or equivalent with at least 7 years of relevant technical experience
- Associate or bachelor's degree in a technical field with at least 5 years of relevant technical experience
- 5+ years of hands-on experience supporting complex electronic systems, enterprise server hardware, integrated data center infrastructure, or comparable production technology environments
- Experience must include diagnosing, repairing, or maintaining enterprise server components, including processors, memory, storage devices, power supplies, BMCs, network adapters, optical connectivity, and copper or fiber cabling
- 3+ years of hands-on experience supporting HPC systems, supercomputing environments, large-scale Linux clusters, or similarly complex Linux-based compute infrastructure
- Experience must include troubleshooting interactions among compute nodes, management systems, high-speed interconnects, storage platforms, operating systems, workload managers or schedulers, power, cooling, and supporting infrastructure
- 3+ years of experience providing technical leadership, mentoring, work coordination, or task direction for a multidisciplinary technical team
- Direct people-management experience is not required
- 3+ years of hands-on Linux system administration, production support, or troubleshooting experience with Red Hat Enterprise Linux (RHEL), SUSE Linux Enterprise Server (SLES), or a comparable enterprise Linux distribution
- Candidates must be able to independently use Linux command-line tools to navigate files
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.