Operations Engineer, Germany
On-siteBerlin, State of Berlin, Germany
Job Summary
Monitor the GTO Operational Dashboard in Datadog, correlate signals across APM, logs, metrics, and RUM to detect anomalies and determine incident ticket necessity. Triage and investigate production issues by creating JIRA tickets, analyzing blast radius, and routing incidents to the appropriate team using the smart routing model. Own lower-severity incidents end-to-end, execute runbooks, and escalate promptly when thresholds are breached or code-level fixes are required. Support the TSO Lead during major incidents by surfacing real-time data, maintaining live incident tickets, and drafting communications for internal updates and status pages. Analyze incident trends and compile data for Post-Incident Review preparation, while building automation scripts and documenting runbooks to improve team efficiency. Conduct structured shift handoffs and cover for the TSO Lead during absences, including severity classification and escalation decisions.
Required Qualifications
- Previous experience working at a gaming company
- 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment
- Strong troubleshooting and investigation skills
- Hands-on experience with Datadog (or equivalent observability platform: Grafana, Splunk, New Relic, Elastic)
- Proficiency in at least one scripting language: Python, Go, or Bash
- Clear written and verbal communication skills in English
- Working knowledge of Kubernetes and cloud infrastructure (GCP preferred, AWS/Azure acceptable)
- Understanding of SLOs, error budgets, and burn-rate alerting
- Experience with incident management tooling: JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence
- Comfort with 24x7 shift-based operations as part of a follow-the-sun model with handoff overlaps
- Weekend on-call (rotating)
- ITIL Foundation certification is a plus but not required
Desired Qualifications
- Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM (Real User Monitoring)
- Experience with distributed systems debugging: tracing failures across microservices, understanding cascading failures, and reading distributed traces end-to-end
- Exposure to database operations (MySQL, PostgreSQL, Redis, Kafka)
- Familiarity with CI/CD pipelines and deployment tooling (GitLab CI, ArgoCD, Helm)
- Experience with JIRA Service Management administration: workflows, automation rules, SLA timers, and queues
- Experience with or strong interest in AI/ML-assisted operations: anomaly detection, alert correlation, predictive monitoring, or automated remediation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.