Sr. Site Reliability Engineer - SRE
$70,000–$115,000 year
RemoteSpain
Job Summary
Design, implement, and maintain highly available, scalable systems while serving as a subject matter expert for observability, including monitoring, alerting, logging, tracing, and synthetic testing. Develop self-service tooling to automate operational tasks, eliminate toil through process improvements, and lead incident response with blameless post-mortems. Define and track SLIs, SLOs, and error budgets, leveraging infrastructure as code and GitOps practices using Terraform, Flux, and GitHub Actions. Provide reliability expertise during system design reviews, document runbooks, and mentor engineers across the organization. Apply AI responsibly to accelerate investigations, improve documentation, and build intelligent operational workflows while maintaining human oversight.
Required Qualifications
- Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering, or Platform Engineering role
- Ability to rapidly build accurate mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability domains
- Strong troubleshooting skills with a methodical, evidence-driven approach to incident response and root cause analysis
- Experience defining and using SLIs, SLOs, and error budgets to guide reliability decisions
- Excellent written and verbal communication skills
- Experience across several of the following areas: Kubernetes platforms, including Amazon EKS, and service mesh technologies such as Istio
- Cloud infrastructure and services within AWS
- Identity and access management systems, including Auth0 and AWS IAM
- Networking fundamentals, including DNS, load balancing, routing, TLS, and connectivity troubleshooting
- GitOps workflows and infrastructure automation using tools such as Flux and Terraform
- Observability platforms and practices, including metrics, logs, traces, alerting, dashboards, and synthetic monitoring
- CI/CD systems and engineering workflows
- Application logging and distributed system debugging
- Demonstrated ability to build and maintain automation, tooling, and self-service capabilities using one or more programming or scripting languages such as Python, Go, or Bash
Desired Qualifications
- Experience defining and working with SLOs, SLIs, and Error Budgets
- Familiarity with other observability tools or concepts beyond Datadog
- Experience with feature flagging platforms like LaunchDarkly
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.