Senior Site Reliability Engineer
$150–$190,000 year
Remote
Job Summary
Operate and maintain all Develocity instances and supporting services, including artifact registries, while participating in a follow-the-sun on-call rotation to troubleshoot incidents across the application and infrastructure stacks. Drive automation for deployment, upgrades, monitoring, self-healing, and recovery, and build observability for logging, metrics, tracing, and alerting. Collaborate with engineering teams to embed reliability into features, own disaster recovery and business continuity, and communicate with customers during incidents and maintenance windows. Optimize performance, resource usage, and costs as we evolve our SaaS operations. Work remotely in a distributed team using asynchronous communication, with a target salary of $150-190k and equity grants.
Required Qualifications
- 5+ years in SRE, DevOps, or equivalent role operating production services at scale
- Strong Kubernetes experience in production environments
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2)
- Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform)
- Track record of incident management and response
- Knowledge of SRE best practices (SLAs, SLOs)
- Scripting proficiency (Python, Bash) for automation
- Experience with 24/7 on-call rotations
- Strong written and verbal English communication
Desired Qualifications
- Experience operating SaaS platforms at scale
- Familiarity with Develocity
- JVM language experience (Java, Kotlin)
- Disaster recovery planning and execution experience
- Customer-facing incident communication skills
- Experience establishing SRE practices in new or growing teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.