Sr. High Performance Computing (HPC) Systems Engineer
$165,000–$230,000 year
On-siteHawthorne, California, United States
Job Summary
Administer and manage HPC clusters, storage systems, and high-speed networks while providing application support to SpaceX personnel across engineering disciplines. Install and integrate Linux-based compute clusters, write instructional documentation, and convey highly technical ideas in non-technical terms. Leverage experience with Kubernetes, cluster resource managers, and container technologies to build, deploy, and troubleshoot large-scale computing environments. Must be eligible for TS/SCI clearance and willing to work extended hours and weekends as needed. Compensation ranges from $165,000 to $230,000 annually, with additional long-term incentives and comprehensive benefits.
Required Qualifications
- Bachelor's degree in computer science, engineering, math, or scientific discipline
- 5+ years of systems engineering experience
- 7+ years of professional experience building software
- 5+ years of hands-on experience with client and server hardware/software, management tools, enterprise networking, virtualization, and security technologies
- Experience with Kubernetes
- Must be willing to work extended hours and weekends as needed
- Eligibility for access to classified material up to TS/SCI with Polygraph
- Must be a U.S. citizen or national
- Must be a U.S. lawful, permanent resident (aka green card holder)
- Must be a Refugee under 8 U.S.C. § 1157
- Must be an Asylee under 8 U.S.C. § 1158
- Must be eligible to obtain the required authorizations from the U.S. Department of State
Desired Qualifications
- 5+ years of professional experience building, deploying and troubleshooting Linux systems
- Experience with a scripting language (Bash, Python) to automate and solve reoccurring tasks
- Experience building, deploying and troubleshooting HPC clusters
- Familiarity with cluster resource managers (Slurm, PBS, LSF)
- Familiarity with monitoring and alerting technologies (Prometheus, Grafana, Nagios)
- Familiarity with scientific and engineering computing (CFD, FEA)
- Familiarity with large scale AI training
- Familiarity with GPU usage in a compute cluster and Cuda
- Experience with containers (Docker, Podman, Singularity)
- Experience deploying and maintaining automated configuration management software (Puppet, Ansible)
- Comfortable working with mission critical and sensitive systems, with a sense of urgency appropriate to the responsibilities
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.