Site Reliability / Production Engineer
RemoteAlgerie Four Corners, Massachusetts, United States
Job Summary
Respond to live incidents by establishing customer impact, investigating system state using logs and metrics, and executing proportionate mitigations such as rollbacks or bounded fixes while validating recovery against real outcomes. Coordinate responses with product teams for deep technical decisions, protect developers from routine pages, and maintain clear ownership of uncertainty and next actions. Improve reliability systems by tuning alerts, strengthening monitoring around customer outcomes, creating runbooks and AI Skills, and automating repeated operational tasks. Work with product teams to close observability gaps and detect unusual service-cost behavior.
Required Qualifications
- Agency and ownership
- Operational judgement
- Technical comfort and aptitude
- AI-native execution
- Accuracy and validation discipline
- Systems thinking
- Clear coordination and communication
- Curiosity and resilience
- Previous responsibility for live production systems or an on-call rota
- The ability to learn an unfamiliar environment, act safely and validate your work
Desired Qualifications
- Experience with every technology in our stack
- A previous SRE job title
- People-management experience
- The ability to recall every command without AI assistance
- Experience with Cloud platforms such as Azure or Cloudflare
- Distributed application and API diagnostics
- Databases, queues and background-processing systems
- Observability, alerting and incident-management platforms
- Infrastructure, deployment and release automation
- Application development and safe production debugging
- AI coding agents and workflow automation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.