Staff Software Engineer, HPC
On-siteSeattle, Washington, United States or Boston, Massachusetts, United States
Job Summary
Design and implement core services and abstractions for distributed compute infrastructure supporting hundreds of thousands of concurrent jobs. Create production-grade APIs, SDKs, and tools that enable engineers to run large-scale distributed workloads across data engineering, AI model training, and simulation. Work with stakeholders to build a multiyear software engineering roadmap, lead cross-team initiatives, and define HPC platform strategy. Optimize job scheduling algorithms, auto-scaling policies, and multi-region orchestration to maximize reliability and resource availability. Identify and resolve systemic reliability and performance issues through profiling and collaboration. Mentor junior engineers while evaluating new technologies to improve computational and storage capabilities.
Required Qualifications
- Experience designing and operating large-scale distributed systems in production
- Experience with Ray.io, particularly Ray Core and Ray Data (or equivalent technologies)
- Experience with Kubernetes, particularly for heterogeneous workloads
- Experience with cloud infrastructure on AWS or similar providers
- Track record of shipping and operating reliable, highly available scalable infrastructure
- Demonstrated ability to prioritize development work and build cross-functional consensus around technical tradeoffs
- Proficiency with Python
Desired Qualifications
- Exposure to machine learning workloads (training, inference, data generation)
- Experience with Kubernetes or SLURM at scale (>10k+ nodes)
- Experience with SLURM workload manager and advanced scheduling policies
- Background in algorithmic optimization or operations research
- Experience building developer tools and platforms used by large engineering organizations
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.