Staff Network Reliability Engineer (Cloud Operations)
RemoteUnited States or Mountain View Santa Clara County, California, United States
Job Summary
Own 24x7 cloud infrastructure health across Skylo's hybrid production environment, managing GKE clusters, on-premise Kubernetes, and persistent storage arrays. Monitor and triage alarms using Prometheus, Grafana, and Loki dashboards to distinguish transient events from systemic degradation. Execute runbooks for P2–P4 faults including node recovery, database failover, and ArgoCD drift remediation without engineering escalation. Serve as the L3 escalation authority for Cloud Infra incidents, diagnosing root causes at the Kubernetes, storage, and database layers. Define and maintain SLOs tied to network SLA commitments, track error budgets, and drive toil reduction through automation partnerships. Lead post-incident analysis, author operational runbooks, and mentor Senior NREs on infrastructure troubleshooting patterns.
Required Qualifications
- 8–10+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment — with direct on-call ownership for Kubernetes-at-scale environments
- Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures
- Hybrid cloud operations: hands-on experience operating both public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms)
- Production observability stack ownership: Prometheus (federation, remote write, WAL management), Grafana, VictoriaMetrics, OpenTelemetry, and alerting pipeline design with Pub/Sub or equivalent
- Database reliability: PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning; Redis cluster operations and persistence management
- GitOps tooling in production: ArgoCD or Flux CD for multi-cluster operations; Helm chart authorship and version management; Terraform or Ansible for infrastructure provisioning
- SRE fundamentals: SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design
- Container and Linux internals: container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding
- Runbook authorship: ability to write infrastructure diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure
- Strong written and verbal communication skills
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.