Boundless Networks logo
Boundless NetworksPosted 3 weeks ago

Senior Infrastructure Engineer - GPU Compute

Remote

Full TimeSenior LevelSmall

Job Summary

Orchestrate a heterogeneous, multi-region GPU fleet using SkyPilot, Kubernetes, and cloud/on-prem providers to schedule inference workloads reliably. Maximize GPU utilization across spot and on-prem capacity while driving down cost per GPU-hour through intelligent workload placement. Tune bare-metal performance via PCIe P2P, NUMA topology, and CUDA driver optimization to push throughput per node. Build secure fleet access with Tailscale and Teleport, alongside robust observability and zero-downtime rollouts for a distributed node fleet. This role requires 5+ years of infrastructure experience and a GitHub profile demonstrating at least one year of activity. Boundless is building the future of AI compute at scale.

Required Qualifications

  • 5+ years of infrastructure/DevOps experience operating large-scale production systems
  • Deep expertise in Kubernetes, Docker, and container orchestration at scale
  • Strong Linux systems administration skills
  • Proficiency in infrastructure-as-code tools (Terraform, Ansible, Pulumi)
  • Track record of managing mission-critical, high-throughput systems
  • Strong infrastructure-as-code background in heterogeneous environments
  • Proficiency in at least one common scripting or programming language (Python, Bash, TypeScript, Go, etc.)
  • Comfort navigating ambiguity with a strong bias for action
  • Candidates must include a public GitHub profile in their application
  • The GitHub profile should demonstrate a minimum of 1 year of activity/history

Desired Qualifications

  • Experience with GPU computing infrastructure (CUDA, bare-metal optimization, kernel tuning)
  • Experience operating ML training or other large-scale distributed compute infrastructure
  • Experience with GPU fleet orchestration (SkyPilot, Ray, Slurm)
  • Familiarity with fleet access and networking tooling (Tailscale, Teleport)
  • Knowledge of network optimization and topology design
  • Experience with multi-region, globally distributed systems
  • Proficiency in Rust or low-level systems programming
  • Experience with on-premises data center operations

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce