Sr Site Reliability Engineer
HybridAustin, Texas, United States
Job Summary
Implement and maintain highly available AWS infrastructure including EKS clusters, Fargate, and multi-region architectures while supporting critical services like Skyway, Frontdoor, and Pantheon. Monitor SLIs, SLOs, and error budgets for Tier 1/2/3 systems, then implement reliability patterns such as circuit breakers and automated failover. Build observability solutions using NewRelic for APM and distributed tracing to reduce MTTD and MTTR, alongside identifying infrastructure cost optimization opportunities through FinOps practices. Execute chaos engineering experiments and participate in game day exercises to identify system weaknesses, while conducting post-incident reviews and supporting on-call rotation for critical systems.
Required Qualifications
- 5+ years in Site Reliability Engineering, DevOps, or Infrastructure Engineering with demonstrated success improving system reliability
- Bachelor's degree or equivalent experience
- 3+ years hands-on experience with AWS (EKS, EC2, RDS, S3, CloudWatch, IAM) and Kubernetes including cluster management
- Proficient programming skills (Python, Go, or Java) with infrastructure automation and Infrastructure as Code experience (Terraform, CloudFormation)
- Production experience with observability tools (NewRelic, Datadog, Prometheus, Grafana, Splunk) and distributed systems
- Experience with CI/CD platforms and GitOps workflows (CircleCI, Argo CD, Jenkins); on-call rotation and incident response
Desired Qualifications
- Exposure to chaos engineering tools
- API Gateway technologies (Tyk/Kong)
- GraphQL federation (Apollo)
- cost optimization initiatives
- FinOps principles
- AWS (EKS, Fargate, Lambda, VPC, Route53, CloudFront)
- Docker
- Istio Service Mesh
- Argo CD
- CircleCI
- Jenkins
- GitHub Actions
- NewRelic - APM, distributed tracing, metrics & logging
- Splunk - logging
- Terraform
- CloudFormation
- Helm
- Kustomize
- Python/Go/Bash
- AWS Secrets Manager
- Vault
- OpsGenie
- PagerDuty
- ServiceNow
- Strong communication skills with ability to explain technical concepts to diverse audiences
- Collaborative approach working across engineering, product, and business teams
- Self-motivated with ability to solve complex problems within established practices and policies
- Data-driven decision making with customer-centric approach and empathy for developer experience
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.