Member of Technical Staff — Accelerator Systems
On-sitePalo Alto, California, United States
Job Summary
Bring up RadixArk's inference and training systems on new accelerator platforms and drive them to competitive performance. Design hardware abstractions that let a single codebase stay fast across vendors without forking, and port and optimize kernels across programming models and memory architectures. Build cross-platform benchmarking, profiling, and regression detection so performance claims hold up on every target, while debugging numerical divergence and correctness gaps between platforms. Work with vendor engineering teams on pre-release hardware, compiler and driver issues, and roadmap feedback to serve as the internal source of truth on platform capabilities. Contribute hardware-specific optimizations, benchmarks, and portability work back to open-source SGLang and Miles.
Required Qualifications
- 4+ years of experience in systems, performance, or ML infrastructure engineering
- Deep expertise in at least one accelerator programming model (CUDA, ROCm/HIP, Pallas/XLA, Triton, or a vendor SDK), with demonstrated ability to pick up new ones quickly
- Strong understanding of accelerator architecture: memory hierarchy, bandwidth limits, occupancy, and the tradeoffs between them
- Experience writing or optimizing high-performance kernels for ML workloads
- Experience with distributed execution and communication libraries (NCCL, RCCL, MPI, or equivalents)
- Proficiency in C++ and Python
- Strong debugging and profiling skills at the system level, including on platforms where the tooling is incomplete or unreliable
- Track record of performance work that shipped into production
Desired Qualifications
- Experience bringing up ML workloads on new silicon
- Hands-on depth in more than one vendor ecosystem
- Experience with compiler stacks (XLA, MLIR, TVM, Triton) or building compiler passes and IR transformations
- Experience designing hardware abstraction layers or portable kernel interfaces
- Quantization and mixed-precision work across differing numeric formats and hardware support levels
- Experience with distributed inference systems (SGLang, vLLM) or training/RL frameworks (Miles, Megatron, veRL, TorchTitan)
- CPU inference optimization (AVX-512/AMX, oneDNN, NUMA-aware execution)
- Experience optimizing collective communication at scale, or scaling workloads to 1000+ accelerators
- Contributions to kernel, compiler, or ML systems open source
- Direct collaboration with silicon vendors or cloud partners on technical evaluations
- Background in HPC or other performance-critical systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.