Senior Infrastructure Engineer - GPU Compute
Remote
Job Summary
Orchestrate a heterogeneous, multi-region GPU fleet using SkyPilot, Kubernetes, and cloud/on-prem providers to schedule inference workloads reliably. Maximize GPU utilization across spot and on-prem capacity while driving down cost per GPU-hour through intelligent workload placement. Tune bare-metal performance via PCIe P2P, NUMA topology, and CUDA driver optimization to push throughput per node. Build secure fleet access with Tailscale and Teleport, alongside robust observability and zero-downtime rollouts for a distributed node fleet. This role requires 5+ years of infrastructure experience and a GitHub profile demonstrating at least one year of activity. Boundless is building the future of AI compute at scale.
Required Qualifications
- 5+ years of infrastructure/DevOps experience operating large-scale production systems
- Deep expertise in Kubernetes, Docker, and container orchestration at scale
- Strong Linux systems administration skills
- Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
- Track record of managing mission-critical, high-throughput systems
- Strong infrastructure-as-code background in heterogeneous environments
- Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
- Comfort navigating ambiguity with a strong bias for action
- Candidates must include a public GitHub profile in their application
- The GitHub profile should demonstrate a minimum of 1 year of activity/history
Desired Qualifications
- Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
- Experience operating ML training or other large-scale distributed compute infrastructure
- Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
- Familiarity with fleet access and networking tooling (Tailscale, Teleport)
- Knowledge of network optimization and topology design
- Experience with multi-region, globally distributed systems
- Proficiency in Rust or low-level systems programming
- Experience with on-premises data center operations
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.