Senior DevOps/SRE Engineer
$140,000–$170,000 year
On-siteChicago, Illinois, United States or Oaks, Pennsylvania, United States
Job Summary
Design and operate cloud infrastructure for critical production applications using multi-account AWS foundations, VPCs, and networking architecture. Build tooling and automation to reduce errors, shorten recovery time, and improve day-to-day operations while implementing reliability guardrails for releases. Manage incident response, perform root cause analysis, and drive durable improvements that prevent recurrence. Apply AI-assisted tools to accelerate infrastructure delivery, troubleshooting, and root-cause analysis, validating outcomes with engineering judgment. Maintain clear, lightweight documentation to support shared ownership and effective on-call operations within an Agile team.
Required Qualifications
- BA/BS, in a related technical field; or the equivalent in education and work experience
- 8+ years of experience in DevOps, SRE, platform engineering, or similar roles supporting application teams running production services
- Strong CI/CD experience (Jenkins and Git-based workflows preferred), including building secure, reliable pipelines and enabling teams to ship safely
- Hands-on, demonstrable experience designing and operating AWS environments, and enabling application teams to adopt AWS correctly (networking, IAM, security, reliability, and cost awareness)
- Infrastructure as Code experience (Terraform preferred; CloudFormation acceptable), including building reusable modules/patterns and managing changes through review and automation
- Experience supporting CI/CD builds and deployment patterns for common application stacks (for example Java, NodeJS, and .NET)
- Experience scripting in Bash, Python, or PowerShell
- Experience working on large scale cloud-based web applications
Desired Qualifications
- Experience implementing and operating observability platforms (logging/metrics/alerting); Elastic Stack/OpenSearch experience is a plus
- Ability to clearly communicate both verbally and in writing with client and team members, including experience documenting and presenting findings
- Excellent analytical skills, organizational abilities, and problem-solving skills
- Familiarity with AI agentic development tools (e.g., Claude Code, GitHub Copilot, Windsurf) and practical experience applying them to infrastructure and operations workflows
- Self-starter who works efficiently in a fast-paced environment with changing priorities and a geographically distributed team
- Ability to think creatively and seek optimum solutions
- Ability to grasp loosely defined concepts and transform them into tangible results and key deliverables
- Diagnostic skills with the ability to analyze technical, business and financial issues and options
- Ability to infer from previous examples, willingness to understand how an application is put together
- Action-oriented, with the ability to quickly deal with change
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.