Site Reliability Engineer
$140,000–$150,000 year
RemoteCalifornia, United States or Arizona, United States
Job Summary
Design and provision AWS infrastructure using Terraform, operate Kubernetes-based workloads, and automate operational processes to ensure system reliability and scalability. Build and improve distributed event-driven systems, define Service Level Indicators and Objectives, and develop automation for deployment, scaling, and incident response. Own platform observability by implementing metrics, logging, and tracing solutions, while leading incident response efforts and facilitating postmortems. Partner with Product and Engineering teams on capacity planning and resilient system design, adhering to HIPAA and SOC 2 compliance requirements. Participate in an on-call rotation supporting production systems.
Required Qualifications
- Three to five (3–5) years of experience in Site Reliability Engineering, DevOps Engineering, Platform Engineering, Cloud Infrastructure Engineering, or similar infrastructure-focused roles
- Bachelor's degree in Computer Science, Information Systems, Software Engineering, or a related technical field
- Equivalent professional experience to the above education requirement
- Strong hands-on experience operating production workloads within AWS environments
- Proven experience managing infrastructure as code using Terraform, including module development, state management, and deployment automation
- Experience operating and supporting production Kubernetes environments
- Hands-on experience deploying and managing applications using Helm
- Experience working with distributed systems, event-driven architectures, or event-sourcing platforms, including concepts such as partitioning, event ordering, replay, and fault tolerance
- Experience establishing and managing observability practices including monitoring, logging, tracing, alerting, and incident response
- Strong understanding of Linux systems administration, networking, cloud architecture, and distributed systems fundamentals
- Experience designing, implementing, and maintaining CI/CD pipelines and deployment automation
- Strong problem-solving skills with the ability to troubleshoot complex infrastructure and application issues
- Excellent written and verbal communication skills with the ability to collaborate effectively across technical and non-technical teams
- High level of ownership, accountability, and initiative with a proactive approach to reliability and operational excellence
- Ability and willingness to participate in an on-call rotation supporting production systems
- Working East Coast business hours (EST)
- Background check, which may include a drug test or other health screenings
- Identity and work eligibility verification using E-Verify upon hire
Desired Qualifications
- Strong programming or scripting experience with Python, Go, or similar languages
- Experience with observability platforms such as Prometheus, Grafana, Datadog, CloudWatch, SigNoz, or OpenTelemetry
- Experience with GitOps tools such as ArgoCD or Flux
- Experience managing databases such as PostgreSQL, MySQL, Redshift, or ClickHouse
- Experience implementing secrets management solutions such as AWS Secrets Manager or HashiCorp Vault
- Experience supporting healthcare technology platforms or other highly regulated environments
- Familiarity with data infrastructure technologies including Snowflake, Redshift, and ETL/ELT pipelines
- Experience with database performance tuning and optimization
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.