Cloud Site Reliability Engineer
RemoteUnited States
United StatesRemoteContractSenior LevelBachelors DegreeEnterprise
ContractSenior LevelBachelors DegreeEnterprise
Job Summary
Design and maintain reliability solutions and SRE utilities to reduce toil and improve cloud platform reliability. Build and optimize Infrastructure as Code using Terraform for AWS resources, develop CI/CD pipelines with automated testing, and define SRE standards including SLIs and SLOs. Apply software engineering best practices to all development, participate in incident management and on-call rotation, and produce clear postmortems. Collaborate within Agile frameworks to deliver cloud automation solutions while staying current with emerging AWS services and SRE methodologies.
Required Qualifications
- Bachelor's degree in computer science, Information Systems, or equivalent background or equivalent experience
- 7+ years of extensive experience in software development with focus on reliability and platform engineering
- 5+ Years of advanced Python development skills with proven experience building enterprise-grade, highly available tools, APIs, and utilities
- 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
- 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
- Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
- Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
- Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
- Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
- Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices
Desired Qualifications
- Experience with GoLang or additional programming languages is a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.