#CORE DIGITAL CAMPUS - IT Operations Specialist High Performance Computing
On-siteAlbacete, Castille-La Mancha, Spain
Job Summary
Monitor the availability, robustness, and performance of HPC clusters while maintaining servers, middleware, networks, and storage in operational condition per OLA/SLAs. Execute support processes for incident, problem, change, configuration, and event management by analyzing error reports, diagnosing issues, and coordinating corrective actions with technical teams and vendors. Coordinate updates to cluster components, LSF scheduler configurations, and deployed predictive solvers. Communicate with stakeholders including end users, Digital teams, vendors, management, and other Airbus divisions. Ensure continuity of operations for over 400 simulation engineers within the Digital Campus ecosystem.
Required Qualifications
- Familiar with Linux server environments
- Familiar with scripting languages (bash, python, etc)
- First experience with HPC systems (as user or administrator)
- Strong interest in deployment, operation and maintenance of complex technical environments
- Good communication skills
- Fluent in English (business & technical), with a collaborative spirit and strong customer orientation
Desired Qualifications
- Linux server administration (RHEL 7/8/9)
- Hardware infrastructure (servers, network switches etc), network storage, network flows & firewall rules
- Ideally experience as HPC administrator: HPC job scheduling (such as LSF or Slurm)
- Knowledge around high-speed interconnect (e.g. Infiniband)
- High Performance Storage (e.g. Panasas)
- HPC node performance engineering
- HPC applications (solver deployment)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.