Senior Site Reliability Engineer
On-siteJakarta, Jakarta, Indonesia
Job Summary
Own availability, performance, scalability, and security of production systems across AWS, GCP, and on-premises environments. Design Kubernetes deployment strategies, manage CI/CD and GitOps pipelines using ArgoCD and Terraform, and diagnose database performance issues. Build observability tools to surface problems proactively and lead structured incident investigations with a metrics-first approach during high-traffic events. Scope and lead medium-to-large infrastructure initiatives by gathering requirements, prioritizing by business impact, and negotiating technical tradeoffs to meet SLAs. Mentor peers and junior engineers while maintaining strict documentation and process discipline.
Required Qualifications
- At least 4 years of experience in either SRE, DevOps, MLOps, or platform engineering
- senior-level scope at a high-traffic company
- Deep expertise in one major cloud provider
- proven ability to ramp up on the other quickly
- Production experience & expertise with Kubernetes
- Linux fundamentals
- CI/CD & GitOps
- ArgoCD or other equivalent stacks
- Database performance analysis & monitoring
- MySQL
- Postgres
- Observability tooling & standards
- Datadog
- OpenTelemetry
- Infrastructure as Code
- Terraform
- Strong working English
- verbal & written communication
- Strong documentation and process discipline
- On-call & incident handling experience during high-traffic events
- Comfortable negotiating with stakeholders to propose technical compromises that meet business SLAs
- Demonstrated growth mindset
- proven ability to own ambiguous scope
Desired Qualifications
- preferably AWS
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.