Storyteller logo
StorytellerPosted 1 month ago

Site Reliability / Production Engineer

Remote

Full TimeSmall

Job Summary

Respond to live incidents by assessing customer impact, investigating system state using logs and metrics, and executing proportionate mitigations such as rollbacks or bounded fixes while validating recovery against real outcomes. Coordinate responses with product teams for deep technical decisions, protect developers from routine pages, and maintain clear ownership of uncertainty and next actions. When quiet, improve reliability by tuning alerts, strengthening monitoring around customer outcomes, creating runbooks, and automating repeated operational tasks. Work within a shared out-of-hours UK coverage rota covering evening shifts and weekday overnight pager duty, with active shifts divided between hires to ensure continuous coverage.

Required Qualifications

  • Agency and ownership
  • Operational judgement
  • Technical comfort and aptitude
  • AI-native execution
  • Accuracy and validation discipline
  • Systems thinking
  • Clear coordination and communication
  • Curiosity and resilience
  • Previous responsibility for live production systems or an on-call rota
  • The ability to learn an unfamiliar environment, act safely and validate your work

Desired Qualifications

  • Experience with every technology in our stack
  • A previous SRE job title
  • People-management experience
  • The ability to recall every command without AI assistance
  • Experience with Cloud platforms such as Azure or Cloudflare
  • Distributed application and API diagnostics
  • Databases, queues and background-processing systems
  • Observability, alerting and incident-management platforms
  • Infrastructure, deployment and release automation
  • Application development and safe production debugging
  • AI coding agents and workflow automation

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce