Systems ML Engineer (Member of the Technical Staff)
On-siteCambridge, Massachusetts, United States
Job Summary
Profile and optimize large-scale model training and inference workloads using tools like Nsight and PyTorch Profiler to identify bottlenecks in data loading and gradient computation. Design high-performance GPU kernels in Triton or CUDA, manage cloud and edge infrastructure for multimodal lab data, and build data pipelines that maximize throughput. Debug performance issues across the full ML stack, from hardware to deployment, while ensuring security, compliance, and safe model updates. Work closely with research and perception teams to squeeze performance out of every shipped model. This role is in-person in Cambridge, MA, with competitive compensation and full benefits.
Required Qualifications
- Experience profiling and optimizing large-scale models for both training and inference
- Comfort writing custom GPU kernels when off-the-shelf ops aren't fast enough
- Experience managing cloud and edge infrastructure that holds up in real lab environments
- Experience with profiling tools (e.g., Nsight, PyTorch Profiler)
- Experience implementing optimizations like kernel fusion, sharding, and tiling
- Experience with PyTorch Distributed
- Experience designing and maintaining high-performance GPU kernels in Triton or CUDA
- Experience designing and optimizing data loading pipelines
- Experience managing deployment across both cloud infrastructure and edge devices
- Experience debugging and resolving performance bottlenecks, resource issues, and failures across the training and deployment stack
- Experience building monitoring, versioning, and rollback into deployments
- Demonstrated expertise in ML systems engineering, including optimizing and deploying large-scale models in production
- Experience debugging and fixing performance and stability issues in deployed systems
- Experience building infrastructure for reproducible, monitored ML deployments
- Experience optimizing inference throughput and resource utilization across cloud and edge
- Deep knowledge of distributed training and serving frameworks, including PyTorch/JAX distributed strategies
- Experience with gradient accumulation, mixed precision training, and checkpoint/recovery systems
- Strong cloud administration skills, including AWS services
- Experience with infrastructure as code (Terraform)
- Experience with Kubernetes orchestration
- Experience with cost optimization, security best practices, and compliance requirements
- Experience deploying and optimizing ML systems on edge or on-prem infrastructure
- Understanding of the ML stack from hardware (GPUs, interconnects, storage) through frameworks (PyTorch, JAX) to deployment and serving
- Skilled at debugging complex failures across the stack – GPU/NCCL issues, data loading bottlenecks, memory leaks, and performance or convergence problems in both training and deployment
- Deep experience optimizing algorithms for cloud and edge environments, including computer vision and other ML algorithms
- Experience with GPU-level work like CUDA and kernel tuning
- Experience working in fast-moving/ambiguous environments (like startups!)
Desired Qualifications
- A passion for and experience in science
- A passion for and experience with AI
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.