Storyteller logo
StorytellerPosted 1 month ago

Site Reliability / Production Engineer

RemoteEgypt

Full TimeSmall

Job Summary

Respond to live incidents by assessing customer impact, investigating system state using logs and AI, and executing proportionate mitigations or rollbacks while validating actual recovery. Coordinate responses by bringing in product teams for deep knowledge, protecting developers from routine pages, and producing clear handovers. Improve reliability by tuning alerts, strengthening monitoring around customer outcomes, creating runbooks and AI Skills, and automating repeated operational tasks. Cover out-of-hours shifts from 17:00 to 01:00 UK time seven days a week, sharing pager duty from 01:00 to 06:00, with rest days and detailed rota arrangements confirmed during hiring.

Required Qualifications

  • Agency and ownership
  • Operational judgement
  • Technical comfort and aptitude
  • AI-native execution
  • Accuracy and validation discipline
  • Systems thinking
  • Clear coordination and communication
  • Curiosity and resilience
  • Previous responsibility for live production systems or an on-call rota
  • The ability to learn an unfamiliar environment, act safely and validate your work

Desired Qualifications

  • Experience with every technology in our stack
  • A previous SRE job title
  • People-management experience
  • The ability to recall every command without AI assistance
  • Experience with Cloud platforms such as Azure or Cloudflare
  • Distributed application and API diagnostics
  • Databases, queues and background-processing systems
  • Observability, alerting and incident-management platforms
  • Infrastructure, deployment and release automation
  • Application development and safe production debugging
  • AI coding agents and workflow automation

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce