Technical Service Operations Lead (TSO Lead), Germany
On-siteBerlin, State of Berlin, Germany
Job Summary
Conduct major incident response as Incident Commander, coordinating cross-functional teams, driving investigations, and managing escalations to meet SLA targets. Draft and send clear updates to leadership, Customer Success, and partners while managing status page communications. Facilitate blameless Post-Incident Reviews, assigning corrective actions and tracking them to closure. Proactively analyze incident trends and production bugs to create Problem tickets and report recommendations to engineering. Enforce the incident management framework, including severity models, priority matrices, and deployment readiness gates. Mentor the Operations Engineer on triage, investigation, and runbook execution while conducting knowledge transfer sessions. Produce shift handoff reports and deliver operational metrics on MTTD, MTTA, MTTR, and proactive detection rates. Audit service catalogue completeness and govern JIRA Service Management workflows. Cover for the Operations Engineer during absences and participate in weekend on-call rotation for critical incidents.
Required Qualifications
- Previous experience working at a gaming company
- 6+ years of experience in incident management, SRE, NOC leadership, or technical operations in a production environment supporting high-availability, high-transaction systems
- Proven incident management experience — coordinating multi-team response, making real-time escalation decisions, and communicating with executive stakeholders under pressure
- Excellent written and verbal communication skills in English — ability to draft clear, concise executive updates at 3 AM under pressure, facilitate blameless PIRs, present operational metrics to senior leadership, and communicate incident status to customers and partners with clarity and professionalism
- Strong ITIL foundation — understanding of incident, problem, and change management lifecycles with practical experience implementing or operating ITIL-aligned workflows
- Technical depth across the observability stack — ability to read and interpret logs, traces, and metrics in Datadog (or equivalent: Grafana, Splunk, New Relic)
- Understanding of APM, SLOs, error budgets, burn-rate alerting, and synthetic monitoring
- Hands-on experience with incident tooling: Datadog, PagerDuty or OpsGenie, JIRA or JIRA Service Management, Slack, and Confluence
- Analytical mindset — ability to identify trends, patterns, and recurring issues from incident data and translate them into actionable recommendations for product and engineering teams
- Experience with SLA/SLO-driven operations where MTTD, MTTA, and MTTR are measured, reported, and improved
- Comfort with 24x7 shift-based operations as part of a follow-the-sun model with handoff overlaps
- Weekend on-call (rotating) for critical severities
Desired Qualifications
- Experience with customer/partner-facing incident communications and status page management
- Experience with or strong interest in AI/ML-assisted operations: anomaly detection, alert correlation, predictive alerting, automated remediation, or self-healing automation
- JIRA Service Management administration experience: workflows, SLA timers, automation rules, queues, and permissions
- Familiarity with Datadog Service Catalog, scorecards, and SLOs — especially burn-rate alerts and multi-window SLOs
- Experience building an operations function from scratch — defining processes, writing runbooks, establishing governance cadences
- Background in Kubernetes, cloud infrastructure (GCP preferred), microservices architecture, or distributed systems
- ITIL certification (Foundation or higher)
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.