Site Reliability Engineer
$165,000–$190,000 year
On-sitePalo Alto, California, United States
Job Summary
Support and maintain the service quality of the customer-facing SaaS security platform by monitoring, debugging, and optimizing production infrastructure on AWS/GCP. Address complex challenges around scalability, reliability, observability, and cost efficiency while collaborating with Engineering teams to maintain Helm charts, application deployment, and CI/CD pipelines. Define service verification strategies within the CI/CD process to meet SLAs and improve developer experience through optimized workflows. Participate in the 24/7 on-call rotation in coordination with the global SRE team.
Required Qualifications
- 3+ years of experience in a DevOps or SRE role supporting SaaS services on GCP and/or AWS
- Bachelor's degree in Computer Science or related field
- Strong proficiency in Kubernetes, microservices architecture, Helm, GitLab CI/CD, and ArgoCD, Prometheus, Grafana
- Programming experience in at least one language
- Deep understanding of autoscaling, version upgrades, and cloud service optimization
- Participate in the on-call rotation, providing 24/7 support in coordination with our global SRE team
- Monitor, debug, and optimize production infrastructure and services on AWS/GCP
Desired Qualifications
- Golang or Python
- Bonus if you're familiar with technologies like Kafka, Elasticsearch, PostgreSQL, ScyllaDB, Databricks, Dagster, Sentry, Kong
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.