SpaceX logo
SpaceXPosted 1 month ago

Site Reliability Engineer — HPC & Automation (Silicon Engineering)

$125,000–$150,000 year

On-siteRedmond, Washington, United States

Full TimeBachelors DegreeEnterpriseAerospace

Job Summary

Deploy, upgrade, operate, and scale high performance computing clusters and services while managing infrastructure as code and modern observability tools. Collaborate with cross-disciplinary teams to develop automated, full turnkey solutions for silicon simulation workflows and operate continuous integration pipelines for building and releasing systems. Identify and eliminate performance bottlenecks using measurement and creative engineering to accelerate chip design iterations and regression turnaround times. Work extended hours and weekends as needed to meet critical milestones for Starlink silicon development.

Required Qualifications

  • Bachelor's degree in computer science, information systems, or an engineering discipline
  • 2+ years of professional experience in system administration, high performance computing, or site reliability engineering
  • 1+ years of development experience with Bash, Python, and/or other programming languages
  • 1+ years of experience with Linux operating systems
  • Ability to work extended hours and weekends as needed to meet critical milestones
  • U.S. citizen or national
  • U.S. lawful, permanent resident (aka green card holder)
  • Refugee under 8 U.S.C. § 1157
  • Asylee under 8 U.S.C. § 1158
  • Eligible to obtain the required authorizations from the U.S. Department of State

Desired Qualifications

  • Familiarity with containerization technologies (i.e. Docker, Kubernetes)
  • Knowledge in computer system concepts (computer architecture, computer organization, operating systems and concurrency)
  • Experience with databases and data modeling (e.g., MySQL, PostgreSQL, SQLite)
  • Networking knowledge of TCP/IP
  • Experience with high performance computing and workload managers (e.g., Slurm, LSF)
  • Experience with Terraform, Ansible, Puppet, or similar automation frameworks
  • Experience building monitoring and alerting as code (e.g., Grafana, Prometheus, custom exporters)
  • Experience with CI/CD automation at scale (e.g., Jenkins, Bamboo, build systems)
  • Experience with infrastructure as code (IaC) tools for managing fleets of servers
  • Experience with using & building REST API clients/servers
  • Experience with enterprise/networked storage automation (e.g., NetApp ONTAP REST API/CLI, NFS)
  • Experience with ASIC design flows and tools (e.g., Cadence, Synopsys, Ansys, Keysight, Siemens)
  • Strong desire to find performance bottlenecks and performance improvement techniques
  • Excellent communication skills with the ability to communicate with customers, peers, management, etc. in both formal and informal situations
  • Ability to quickly learn new tools and frameworks
  • Interest in or experience with AI/LLM-assisted tooling (e.g., Grok, Claude Code)

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce