Subject Matter Data Center Network Engineer
On-siteAlbuquerque, New Mexico, United States
Job Summary
Provide technical analysis and design for specialized applications in mission-critical production environments. Perform functional systems analysis, integration, and documentation for complex engineering and scientific systems. Apply advanced principles to solve technical problems and develop automated solutions for networking, telecommunications, and automation. Analyze user needs to define requirements and develop plans for moderately complex to extremely complex systems. Support planned change windows and adhere to structured rollback procedures while maintaining U.S. Department of Energy clearance. Participate in on-call rotations and after-hours support as needed. This role supports a major national laboratory within a government contracting firm specializing in HPC networks and data center infrastructure.
Required Qualifications
- BS in relevant discipline
- 3 years, or more, of directly related experience
- Ability to obtain & maintain a U.S. Dept. of Energy Clearance
- U.S. Citizenship
- Experience working in mission‐critical production environments with change control, incident response, and structured troubleshooting
- Ability to work safely and effectively in live data center spaces (DCFIT coordination, raised floor environments, cabling standards)
- Strong analytical and troubleshooting abilities, with a demonstrated ability to drive issues to resolution during outages
- Clear verbal and written communication skills
- Ability to participate in on‐call rotations and support after‐hours change windows if needed
- Strong understanding of L2/L3 networking fundamentals: VLANs, STP, LACP, static routing, OSPF, BGP, ACLs, QoS
- Experience configuring and supporting major network vendors (any of): Cisco, Arista, Juniper, Mellanox/NVIDIA Networking
- Familiarity with multipath topologies, spine‐leaf architectures, and high‐bandwidth fabrics
- Understanding of basic firewall and segmentation concepts (paths, zones, NAT, routing symmetry)
- Familiarity with enterprise cabling best practices: fiber types (LR/SR/ER), transceiver selection, rack elevation awareness
- Working knowledge of monitoring and telemetry: CloudVision, Netropy/Apposite tools, Nagios/Prometheus/Grafana, or similar
- Ability to follow SOPs/MOPs, support planned change windows, and adhere to structured rollback/validation procedures
- In lieu of degree, 9 years of Related experience may be substituted for relevant education and vice versa
Desired Qualifications
- Hands‐on experience with Arista EOS, Cisco NX‐OS, Juniper JunOS, or Mellanox/NVIDIA platforms supporting 40/100/400Gbps networks
- Experience with BGP tuning, traffic‐engineering policies, ECMP management, asymmetric‐routing detection, and complex route‐map design
- Familiarity with EVPN/VXLAN, modern DC underlays, and high‐availability routing gateway designs (HA router pairs, blue/green cutovers)
- Experience troubleshooting firewall pathing issues, zone interactions, NAT64/NAT policy behavior, and segmentation used for HPC
- Exposure to HPC networks and interconnects: RDMA, RoCE, basic InfiniBand concepts (subnet managers, fabric behavior)
- Understanding of congestion behaviors common in HPC (parallel I/O bursts, GPU‐node communication patterns)
- Familiarity moving or supporting HPC data flows across multi‐site infrastructures (tri‐lab WAN, DisCom, IHPC routing)
- Understanding of distributed HPC storage systems (Lustre, BeeGFS, Isilon, Spectrum Scale) and the routing or MTU constraints around them
- Ability to work with HPC teams during large cluster deployments (rack placement coordination, switch firmware updates, cabling checks)
- Hands‐on experience with automation or configuration management tools (Ansible, Python, Terraform)
- Ability to build or maintain dashboards and telemetry for HPC network utilization (CloudVision pipelines, Netropy test data)
- Familiarity with gNMI/gRPC or streaming telemetry for performance analysis and anomaly detection
- Understanding of high‐availability designs, reduced failure domains, and deterministic failover patterns in HPC networks
- Experience monitoring and optimizing latency, packet loss, MTU mismatches, and congestion across crypto tunnels or 100Gbps transport
- Ability to assist with capacity planning for rapidly growing HPC systems (rack density, port utilization, fiber paths, WAN circuit expansions)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.