Production Engineer, Site Reliability (Application Software)
$125,000–$160,000 year
On-siteHawthorne, California, United States
Job Summary
Deploy, upgrade, operate, and scale mission-critical products and services while managing infrastructure as code with modern observability tools. Collaborate with software engineers to design highly operable systems and engage in the full software development lifecycle from inception to continuous refinement. Practice sustainable incident response, conduct blameless postmortems, and provide high-quality end-user support to vehicle software engineers. Identify and eliminate performance bottlenecks using measurement and creative engineering while participating in the team's on-call rotation. This role significantly reduces safety-critical build and test times for Starship vehicle software, supporting Falcon 9, Starship, and Dragon missions alongside Starlink's global growth.
Required Qualifications
- Bachelor's degree in computer science, information systems, or an engineering discipline
- 3+ years of professional experience in SRE or DevOps in lieu of a degree
- 1+ years of experience with Python and Python-based development frameworks
- Experience with Linux operating systems
- Must be able to work extended hours and weekends as needed
- Must be a U.S. citizen or national
- Must be a U.S. lawful, permanent resident (aka green card holder)
- Must be a Refugee under 8 U.S.C. § 1157
- Must be an Asylee under 8 U.S.C. § 1158
- Must be eligible to obtain the required authorizations from the U.S. Department of State
Desired Qualifications
- Experience with build systems (Bazel, Buck, Make, etc.)
- Experience with both container and virtualization technologies (Docker, Kubernetes, vSphere, QEMU, KVM, etc.)
- Experience with databases and data modeling (Postgres, MySQL, ClickHouse, etc.)
- Experience with infrastructure as code (IaC) tools for managing fleets of servers
- Experience with Terraform, Ansible, Puppet, or similar automation frameworks
- Knowledge of the technologies that predate and underpin modern cloud infrastructure, with the ability to translate high-level developer experiences into specific implementations from first principles
- Ability to work with mission-critical and sensitive systems with appropriate urgency and care
- Ability to communicate effectively with customers, peers, and management in both formal and informal settings
- Experience with full-stack development (the team primarily uses Python, JavaScript, and C#; end users primarily use C++)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.