Senior Site Reliability Engineer (all genders)
HybridBerlin, State of Berlin, Germany
Job Summary
Define SLOs, SLIs, and Error Budgets across Next Generation and Infinity products to drive data-based reliability decisions. Lead incident response by enabling rapid detection, clear communication, and blameless postmortems while implementing follow-ups. Reduce manual toil through automation and GitOps practices using Argo CD and Flux, building self-healing and self-service capabilities. Support the NG Search Operator development and introduce auto-scaling primitives like HPA, VPA, KEDA, and Cluster Autoscaler. Enhance observability by developing metrics, logs, traces, alerting, and runbooks that support on-call operations. Plan capacity and costs across on-premise sites in Frankfurt and Stockholm, alongside public cloud burst scenarios, leveraging AI tools for diagnosis and workflow optimization.
Required Qualifications
- Erfahrung als SRE, Infrastructure oder Production Engineer in einem SaaS- oder Plattform-Umfeld
- starker Software-/Operations-Hintergrund mit klarem Willen, in die SRE-Rolle hineinzuwachsen
- Solides Verständnis von SLOs, Error Budgets, Incident Management und Observability
- Hands-on Erfahrung mit Kubernetes
- Interesse an Cluster-Lifecycle, Upgrades und Operator-Pattern
- Erfahrung oder starkes Interesse an Harvester bzw. vergleichbaren HCI-/Virtualisierungsplattformen
- Kenntnisse zu Auto-Scaling-Primitiven (HPA, VPA, Cluster Autoscaler, KEDA) und Kapazitätsplanung on-prem und in der Cloud
- Verständnis für Netzwerke in produktionsnahen Rechenzentren
- Ausgeprägter Automatisierungsinstinkt und eine Haltung, Toil strukturell zu eliminieren
- Praktische Erfahrung im Einsatz von AI-Tools im operativen Betrieb
- Sehr gute Englischkenntnisse
- Deutsch von Vorteil
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.