Senior Site Reliability Engineer (SRE)
$175–$229,000 year
HybridPalo Alto, California, United States
Job Summary
Build and operate the manufacturing acceleration platform that captures digital exhaust from assembly lines to identify complex engineering insights. Deploy and maintain commercial SaaS platforms on AWS, leveraging Linux, Kubernetes, Terraform, and AI tools to optimize manufacturing yield and throughput. Drive impactful projects independently, automate repetitive infrastructure tasks, and measure KPIs to ensure system reliability and efficiency. This role requires 5+ years of DevOps/SRE experience and comfort with high-growth ambiguity. Candidates must be U.S. citizens due to government contract requirements.
Required Qualifications
- 5 or more years of DevOps or SRE experience deploying and operating commercial SaaS platforms on public cloud infrastructure, AWS preferred
- Expert knowledge with Linux, shell, containerization, Kubernetes, IaC (terraform preferred), monitoring, logging, and APM tools
- Proven ability to take initiative and drive impactful projects to completion efficiently and independently
- Comfort with ambiguity, pace, and frequent pivots inherent in a startup environment, with a track record of creating clarity for teams
- Experience introducing and integrating AI tools/processes into development and operation workflows
- Demonstrated skill in setting, iterating on, and measuring KPIs to ensure ongoing performance, reliability and efficiency
- Dead serious about performance, scalability, and reliability (PSR)
- Systems engineering & infrastructure expertise
- Automation, automation, automation
- Operating in ambiguity & high-growth environments
- Dependable, trustworthy
Desired Qualifications
- Network/application security and compliance experience is a plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.