Helix AI Engineer, Training Performance
$200,000–$400,000 year
On-siteSan Jose, California, United States
Job Summary
Optimize training performance for 100B+ parameter models across 100k+ GPUs by writing custom Triton/CUDA kernels and extending kernel compilers like Triton and Gluon. Collaborate on accelerator selection, cluster topology, and scheduling while building dashboards for continuous performance monitoring and root-cause analysis. Improve data loading pipelines, checkpointing, and fault tolerance to ensure large jobs recover quickly from node failures. Partner with researchers to co-design model architectures and training recipes that maximize hardware utilization, and evaluate emerging accelerators to lead proof-of-concept ports. This role supports Figure's goal of shipping autonomous humanoids with human-level intelligence by pushing distributed training frameworks to scale.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field
- 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
- Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
- Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
- Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement)
- Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
- Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
- Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities
Desired Qualifications
- Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
- Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
- Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.