Designworks Talent logo
Designworks TalentPosted 1 month ago

AI Training Infrastructure Engineer

$150,000–$250,000 year

HybridBellevue, Washington, United States

Full TimeSenior LevelStartup

Job Summary

Build and scale distributed training infrastructure supporting large AI models across GPU clusters. Design systems that increase training reliability, efficiency, and resource utilization while developing solutions for fault tolerance, checkpointing, and recovery. Integrate models into production pipelines and diagnose issues impacting throughput and cost. Establish best practices for operational processes and developer experience tools. Collaborate with infrastructure and machine learning teams to solve complex challenges in distributed computing and production readiness.

Required Qualifications

  • Hands-on experience building and operating distributed training systems or large-scale machine learning infrastructure
  • Experience supporting large AI models, foundation models, post-training workflows, or similar ML systems
  • Strong understanding of the reliability, scalability, and efficiency challenges associated with multi-node GPU training
  • Experience integrating training systems with production machine learning pipelines
  • Strong programming skills and experience working with complex distributed systems
  • Ability to independently own technically challenging projects in a fast-moving engineering environment
  • Comfortable operating with high ownership and limited process overhead
  • U.S. work authorization
  • Candidates elsewhere in the U.S. who are open to relocation

Desired Qualifications

  • Experience with distributed training frameworks such as PyTorch Distributed, DeepSpeed, Megatron-LM, Ray, or similar technologies
  • Experience with supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), or other post-training workflows
  • Background operating AI training infrastructure at scale within a hyperscaler, AI research organization, cloud provider, or GPU cloud environment
  • Experience optimizing GPU utilization, training performance, or distributed system reliability
  • Familiarity with Kubernetes, containerized AI workloads, and large-scale infrastructure platforms

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce