ML Infrastructure Engineer
$180,000–$440,000 year
On-sitePalo Alto, California, United States
Job Summary
Design and scale GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid ML hypothesis iteration. Develop data pipelines integrating large-scale training and inference systems while collaborating with ML teams to productionize models and ensure seamless stack integration. Guarantee scalability, reliability, and efficiency for large-scale machine learning systems by working across the full stack to solve complex problems independently. Mentor junior engineers and contribute to team growth. This role supports xAI's mission to create AI systems that understand the universe, requiring strong prioritization, communication, and hands-on engineering excellence within a flat organizational structure.
Required Qualifications
- Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
- 2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
- 2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
- Strong proficiency with Python and experience with compiled languages such as C++ or Rust
- All employees are expected to have strong communication skills
- They should be able to concisely and accurately share knowledge with their teammates
- Work ethic and strong prioritization skills are important
- Leadership is given to those who show initiative and consistently deliver excellence
Desired Qualifications
- Deep familiarity with modern ML frameworks such as JAX or PyTorch
- Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
- Comfortable with Linux systems and orchestration tools
- Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.