Nimble Robotics logo
Nimble RoboticsPosted 1 week ago

Senior/Staff Software Engineer, Infrastructure (ML)

$210,000–$210,000 year

On-siteSan Francisco, California, United States

Full TimeSenior LevelSmall

Job Summary

Design and maintain ML training infrastructure that enables the AI team to run training jobs efficiently and manage experiments quickly. Build low-latency inference pipelines for production robotics workloads and develop, tune, and optimize low-level CUDA kernels. Design training-platform systems for scalable model training, including high-throughput data ingestion, dataset sharding, and sampling for distributed training. Participate in design reviews to evaluate technical tradeoffs and review code to uphold best practices around style, correctness, and performance. Mentor junior engineers and contribute to documentation as systems evolve. Own end-to-end GPU utilization and make runs reproducible so researchers can launch experiments with a single command.

Required Qualifications

  • Bachelor's, Master's, or PhD in Computer Science or a related field, or equivalent practical experience
  • 4+ years of industry experience in infrastructure, distributed systems, ML systems, robotics, or a related area
  • Experience with programming languages such as Rust, Go, Python, or C++
  • Experience with ML frameworks such as PyTorch or JAX
  • Strong understanding of distributed systems, systems programming fundamentals, memory management, and performance optimization
  • Experience with Kubernetes orchestration, resource scheduling for large distributed jobs, and containerized deployment pipelines
  • Ability to debug and optimize bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations
  • Ability to reason from first principles and optimize systems for both memory-bound and compute-bound workloads
  • Strong cross-functional communication skills, ownership, and a growth mindset

Desired Qualifications

  • Hands-on experience with distributed training frameworks and techniques such as PyTorch DDP/FSDP, DeepSpeed, Megatron, or NCCL
  • Hands-on experience with GPU kernel development
  • Experience with data engineering technologies such as Parquet, Arrow, or similar systems

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce