Virtualization & Orchestration Engineer
$150,000–$250,000 year
HybridBellevue, Washington, United States
Job Summary
Design and build virtualization infrastructure supporting GPU-intensive AI and HPC workloads. Develop and operate Kubernetes-based orchestration systems for cluster provisioning and workload scheduling. Build automated provisioning systems that enable GPU capacity to be allocated, scaled, and reclaimed efficiently across multiple tenants. Design solutions for workload placement, resource management, and cluster lifecycle operations. Partner closely with hardware, networking, and AI platform teams to ensure orchestration systems align with real-world constraints. Improve the reliability, security, and operational maturity of the platform while contributing to architectural decisions and engineering standards. This foundational role within the company's largest infrastructure organization focuses on designing and operating systems that make GPU capacity available, scalable, secure, and reliable across a multi-tenant AI cloud platform.
Required Qualifications
- Strong hands-on experience with Kubernetes and container orchestration in production environments
- Experience designing, building, and operating large-scale infrastructure platforms
- Background with virtualization technologies supporting cloud, HPC, GPU, or distributed computing environments
- Understanding of GPU cluster provisioning, workload scheduling, and resource management
- Experience with Linux-based infrastructure and distributed systems concepts
- Ability to independently own complex systems from design through production operation
- Comfortable working in a fast-moving environment where architecture and processes are being established
- U.S. work authorization
- Hybrid role based in the Bellevue, WA area
- Approximately three days per week in the office
- Candidates elsewhere in the U.S. who are open to relocation
Desired Qualifications
- Experience with GPU scheduling technologies such as Slurm, Kubernetes device plugins, NVIDIA GPU Operator, or similar frameworks
- Experience supporting AI infrastructure, machine learning platforms, HPC environments, or GPU cloud providers
- Background building multi-tenant infrastructure platforms for cloud providers or large-scale compute environments
- Experience with infrastructure automation, Infrastructure as Code, and platform engineering practices
- Familiarity with high-performance networking and GPU cluster architectures
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.