MLOps Engineer - GPU Platform (Inference & Training)
$150,000–$200,000 year
On-siteAlmaty, Almaty, Kazakhstan
Job Summary
Optimize distributed training clusters by tuning NCCL, managing InfiniBand/RoCE fabric health, and implementing topology-aware scheduling to maximize GPU utilization. Own the Talos/Sidero Omni cluster lifecycle across multi-provider fleets, handling node bootstraps, driver upgrades, and incident escalation. Operate inference autoscaling via KEDA and Kafka queues, while maintaining GitOps infrastructure with ArgoCD and Terraform for all cluster changes. Monitor fleet health through VictoriaMetrics and Prometheus dashboards, closing loops on idle allocations and capacity mismatches identified by AIOps tools. Partner with ML engineers on model rollouts and runtime configurations to ensure training throughput and inference efficiency.
Required Qualifications
- 3+ years running production Kubernetes as SRE/Platform/MLOps, including GPU workloads
- Hands-on distributed training operations: NCCL, high-speed interconnects (InfiniBand/RoCE), multi-node job scheduling, checkpointing strategies
- Bare-metal Kubernetes experience: Talos or similar immutable-OS setups; node lifecycle without a cloud safety net
- The NVIDIA stack: drivers, container toolkit, GPU Operator, DCGM metrics
- GitOps fluency (ArgoCD/Flux + Helm) and Terraform
- strong Linux and networking (multi-cluster, VPN/TGW topologies)
- Queue-based autoscaling (KEDA/HPA) and enough Kafka to reason about consumer lag
- Python or Go for automation
- You will own the GPU platform behind our generative video/image products - both the inference fleet that serves production traffic and the training clusters where our models are built
- The estate spans multiple providers orchestrated with Kubernetes on Talos Linux managed by Sidero Omni
- The fleet target is >95% GPU utilization
- Optimize the training clusters: distributed training at scale - NCCL tuning, InfiniBand/RoCE fabric health, topology-aware scheduling and gang placement, GPU/network throughput, fast checkpointing, job preemption and recovery
- Own Talos / Sidero Omni cluster lifecycle across the GPU fleet: node bootstrap and upgrades, GPU drivers / NVIDIA GPU Operator / DCGM on an immutable OS, zero-downtime rollouts
- Operate the multi-provider GPU fleet: capacity planning across Nebius regions and bare-metal RTX Pro pools, hardware incident escalation to providers, node lifecycle (NotReady triage, XID errors, driver upgrades)
- Own inference autoscaling: KEDA-driven, Kafka-queue-based scaling of GPU consumers; GPU-aware scheduling; warm pools and cold-start reduction; supply/demand tuning of our in-house autoscaler (higgscaler)
- GitOps everything: ArgoCD multi-cluster (10+ clusters from one repo), Helm, Terraform (HCP)
- Observability & SLOs: VictoriaMetrics/Logs/Traces, Prometheus, DCGM exporters
- GPU efficiency as a discipline: hunt idle allocations, capacity/demand mismatches, starved queues
- Partner with ML engineers on training runs and model-serving rollouts (runtimes, batching, memory sizing) and with the core team on AWS EKS (Karpenter, Istio, Bottlerocket, gVisor sandboxes)
- Must be able to debug 'GPU visible but not allocatable' at 3am
- You have the habit of measuring throughput before and after every change
- On-site role in our Almaty office (we will relocate you from anywhere)
Desired Qualifications
- C++ and CUDA programming: custom kernels, memory/occupancy tuning, profiling with Nsight Systems/Compute
- Sidero Omni in production; multi-provider GPU clouds (Nebius, CoreWeave, Lambda)
- Inference runtimes (Triton, vLLM, TensorRT) and batching economics
- FinOps for GPU fleets: cost per generation, commitment planning
- Experience building internal platforms or AIOps tooling
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.