Network Automation & Reliability Engineer
$156,000–$176,800 year
RemoteUnited States
Job Summary
Design and maintain Python tooling that provisions, validates, and audits network devices across a global fleet of data centers and POP sites. Build model-driven provisioning libraries using Jinja2 templating with NETCONF/YANG or REST APIs to codify golden configurations and eliminate drift. Develop automated remediation for recurring failure classes triggered by syslog and telemetry, while reducing operational toil through tested, reviewed code. Operate and scale data center fabrics, backbone links, and out-of-band management networks, designing and tuning BGP policy to control path selection and eliminate single-carrier points of failure. Execute zero-downtime change management, including hitless migrations and staged rollbacks, while supporting capacity expansion and site turn-ups. Operationalize multi-vendor streaming telemetry (gNMI/gRPC, OpenConfig) and build observability for hop-by-hop path tracing and fast root cause analysis. Participate in a 24x7 on-call rotation to lead incident response, write RCAs, and drive follow-up automation. Maintain firewall and ACL policies across multi-vendor platforms and support SIRT/PSIRT CVE remediation through automated regression testing.
Required Qualifications
- Python — primary requirement: Demonstrated, sustained Python development in a production network or infrastructure environment. You have written and maintained tooling that other engineers depended on: config generation and validation, API integrations, telemetry collectors, automated remediation, or test harnesses. You are comfortable with modules, packaging, testing, code review, and version control — not just single-file scripts.
- Routing & switching depth: Production experience with BGP (policy, path selection, multihoming), plus IS-IS or OSPF, ECMP, and VXLAN/EVPN or MPLS overlays.
- Multi-vendor hardware: Hands-on operations across at least two of Arista EOS, Juniper Junos (QFX/SRX/PTX/MX), or Cisco IOS-XR / NX-OS.
- Automation frameworks: Ansible and Jinja2, plus NETCONF/YANG, RESTCONF, or vendor REST APIs for model-driven configuration management.
- Telemetry & monitoring: gNMI/gRPC streaming telemetry, OpenConfig models, SNMP, flow telemetry, and dashboarding/alerting on top of them.
- Linux: Comfortable operating on Linux hosts — networking stack, packet capture, systemd services, and shell scripting.
- Version control and CI: Git-based workflows with peer review; experience shipping network changes through a pipeline (Jenkins, GitLab CI, or GitHub Actions).
- Production on-call: Experience holding a 24x7 rotation for a live network, including incident command and root cause analysis.
- Experience level: Roughly 2–5 years in a network production, network reliability, or network automation role.
- Location & work authorization: Must be located in the United States and authorized to work in the US.
Desired Qualifications
- Python you actually wrote, described as software, not as a skills keyword. Name the project, what it did, roughly how large it was, and who used it. 'Python (scripting)' in a skills list will not clear this bar.
- A GitHub, GitLab, or public repo link is strongly preferred— we look at code.
- The specific Python network libraries you have used in production — for example Netmiko, NAPALM, Nornir, pyATS/Genie, Scrapli, ncclient, or a vendor SDK — and what you built with each.
- BGP in production at scale. Name the policy work: local preference, MED, communities, import/export policy, route reflection, ECMP, or multihoming across carriers.
- Hands-on with at least two of Arista EOS, Juniper Junos (QFX / SRX / PTX / MX), or Cisco IOS-XR / NX-OS — named by platform, not just by vendor.
- Telemetry and observability you implemented: gNMI/gRPC, OpenConfig, streaming telemetry, flow telemetry, or SNMP-to-TSDB pipelines. Say what you subscribed to and what you did with the data.
- Config-as-code: Jinja2 templating, NETCONF/YANG or REST-API-driven provisioning, golden configs, ZTP, and the Git/CI workflow you shipped changes through.
- Your networking certifications, named, with dates and credential IDs or verification links— we verify certifications.
- Whether you have carried production on-call, and at what scale (sites, devices, or POPs).
- Shortly after you apply you will receive a short role-specific questionnaire — completing it promptly is the fastest way to move into screening.
- Out-of-band network experience — console server fleets (ZPE Nodegrid, OpenGear), RS-232 configuration management, or OOB build-out for new data center capacity.
- Data center or POP build-out and turn-up: new product introduction (NPI) for switching platforms, port channelization and optics selection, cabling and rack density planning, or hardware qualification and stress testing.
- MACsec, IPsec at scale, or zero-trust segmentation with 802.1X and NAC.
- Optical or transport exposure — DWDM, Ciena, PON/OLT/ONU, or IXIA/Spirent test automation.
- Building or contributing to a network source of truth (NetBox or in-house) aggregating BGP, link-state, and drain-state data.
- Applying LLM or agentic tooling to on-call workflows — automated triage, runbook execution, or incident summarization.
- A master's degree in Network Engineering, Telecommunications, or Computer Science.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.