Figure logo
FigurePosted 1 week ago

Helix AI Engineer, Training Performance

$200,000–$400,000 year

On-siteSan Jose, California, United States

Full TimeMasters DegreeSmallAI Robotics

Job Summary

Optimize training performance for 100B+ parameter models across 100k+ GPUs by writing custom Triton/CUDA kernels and extending kernel compilers like Triton and Gluon. Collaborate on accelerator selection, cluster topology, and scheduling while building dashboards for continuous performance monitoring and root-cause analysis. Improve data loading pipelines, checkpointing, and fault tolerance to ensure large jobs recover quickly from node failures. Partner with researchers to co-design model architectures and training recipes that maximize hardware utilization, and evaluate emerging accelerators to lead proof-of-concept ports. This role supports Figure's goal of shipping autonomous humanoids with human-level intelligence by pushing distributed training frameworks to scale.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field
  • 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
  • Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
  • Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
  • Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement)
  • Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
  • Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
  • Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities

Desired Qualifications

  • Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
  • Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
  • Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce