Senior Lead Network Engineer
On-siteCairo, Cairo, Egypt
Job Summary
Develop network configurations and architectures for Ethernet and InfiniBand fabrics in a high-performance computing and AI environment. Operate, maintain, and support these networks while performing ongoing upgrades, lifecycle management, and hardware diagnostics. Monitor health, performance, and capacity to ensure reliable, low-latency data flow, responding to and resolving incidents within the client ticketing system to meet SLAs. Build automation and tooling to improve operational efficiency and lead design reviews for knowledge sharing. Work within a weekly on-call rotation to address after-hours infrastructure issues.
Required Qualifications
- 14+ years of hands-on experience supporting enterprise or data center-scale networks
- Experience working in HPC, AI/ML, or performance-sensitive environments
- Practical experience administering InfiniBand (Mellanox/NVIDIA) and Ethernet (Cumulus, SONiC) networks
- Strong understanding of data center networking concepts, including servers, storage, and high-speed interconnects
- Solid knowledge of Layer 2 and Layer 3 networking, including routing and switching fundamentals
- Installing, monitoring, and maintaining very large-scale data center networks
- Low-latency, high-bandwidth fabric support and performance tuning for distributed compute and GPU workloads
- VXLAN/EVPN architectures and routing protocols such as BGP and OSPF
- Exposure to communication libraries such as NCCL, UCX, and MPI
- Network management and monitoring tools: UFM, OpenSM, NetQ, or similar
- Ability to troubleshoot and resolve network issues in complex, distributed environments
- Strong documentation and communication skills
- Proven ability to work effectively as part of a team and provide operational support
- Participate in a weekly on-call rotation and respond to network and infrastructure issues after hours when required
Desired Qualifications
- Hands-on experience with high-performance / parallel storage environments (e.g., Lustre, GPFS/Spectrum Scale, BeeGFS, Ceph, NVMe-oF), including storage networking and I/O performance troubleshooting
- Production Linux systems administration at scale — provisioning, configuration management (Ansible/Salt), kernel/network stack tuning, schedulers (Slurm), containerization
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.