Sciforium logo
SciforiumPosted 2 weeks ago

GPU Cluster Engineer, Networking

$150,000–$180,000 year

On-siteSan Francisco, California, United States

Full TimeSenior LevelStartup

Job Summary

Design the complete networking stack for large-scale GPU clusters, from RDMA compute fabrics to cloud connectivity. Architect greenfield topologies including InfiniBand and RoCE v2, specifying switches, optics, and cabling plans. Configure and validate lossless Ethernet, subnet managers, and BGP/EVPN-VXLAN underlays while benchmarking performance with perftest and nccl-tests. Maintain production routing, firewalls, and security domains using network-as-code workflows with Ansible and Git. Operate multi-vendor NOSes, manage firmware upgrades, and serve as the escalation point for fabric incidents. Support cross-site dark fiber links and hybrid cloud interconnects to ensure the network never becomes the bottleneck for frontier AI models.

Required Qualifications

  • 7+ years designing and operating production data center networks, including at least one large-scale HPC/AI cluster fabric (InfiniBand or RoCE v2) you designed or ran.
  • Expert-level routing and switching: BGP (including EVPN-VXLAN), OSPF, ECMP, VRF segmentation, across at least two major vendor platforms.
  • Deep RDMA expertise: lossless RoCE v2 tuning (PFC/ECN/DCQCN) and/or InfiniBand fabric management (subnet managers, adaptive routing, SHARP).
  • Hands-on experience with 100–800G optics, high-radix switching, and Clos/rail-optimized topologies.
  • Production network security experience: enterprise firewalls, site-to-site and remote-access VPNs, segmentation design.
  • Network automation proficiency: Python plus Ansible (or Nornir/NAPALM), with Git-based configuration workflows.
  • Working knowledge of the host-side RDMA stack (MOFED/DOCA, NIC tuning, GPUDirect) and how NCCL/RCCL collectives map onto physical fabric.
  • Cloud networking experience: VPC design and dedicated interconnects (Direct Connect/ExpressRoute class).

Desired Qualifications

  • Dark fiber / DWDM procurement and operations experience.
  • BlueField DPU or SmartNIC deployments.
  • Experience supporting distributed training at 1,000+ GPU scale, or multi-cluster/cross-site training.
  • Kubernetes networking for GPU serving (CNI, SR-IOV, Multus).
  • Expert certifications (CCIE/JNCIE) or NVIDIA networking certification (InfiniBand/Spectrum-X/UFM).

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce