Site Reliability Engineer (Network)
On-siteSan Francisco, California, United States or Golden, Colorado, United States
Job Summary
Design and implement Loft's network architecture, including cloud networking, VPN/site-to-site segmentation, and routing across offices, ground stations, and test equipment. Tighten the network security posture through segmentation, access control, and connectivity hardening while managing network infrastructure as code via IaC provisioning, CI/CD, and GitOps workflows. Define SLOs for connectivity and ground-station availability, own observability using metrics, logs, and tracing, and investigate reliability issues with root-cause analysis to reduce operational toil. Contribute to the SRE team's broader incident-response work and foster the SatDevOps culture. This role offers a rotating Flight Director opportunity to manage the health and safety of the satellite fleet, ensuring reliable global connectivity for end-to-end mission operations.
Required Qualifications
- 4–5 years in network engineering, with deep hands-on networking: routing, VPN/site-to-site, segmentation, DNS, firewalling
- Hands-on Software-Defined Networking: comfortable managing networks programmatically / as code rather than appliance-by-appliance (everything we run is IaC)
- Strong public-cloud networking experience, ideally GCP (VPCs, peering, hybrid connectivity)
- Comfortable operating on a Kuburnetes platform (k8s, Docker) and occasionally writing code
- Infrastructure-as-Code (Terraform or similar) and a CI/CD-driven, GitOps workflow
- An SRE mindset: SLOs, observability, and reducing toil through automation
- Ability to debug connectivity issues in a complex, hybrid network environment
- Degree in Computer Science or a related field, or equivalent practical experience
Desired Qualifications
- Network certifications (CCNP or similar)
- Hands-on experience with GitOps frameworks (ArgoCD, FluxCD)
- Interest or experience in FinOps and cost-optimized architectures
- Familiarity with security practices: vulnerability scanning, threat detection, risk mitigation
- Understanding of orchestration in resource-constrained environments, like space systems
- Grafana-centric stacks a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.