Site Reliability Engineer (SRE)
$163,710–$306,000 year
HybridNew York City, New York, United States or San Francisco, California, United States
Job Summary
Own reliability across Retool Cloud, managed single tenant, BYOC, and self-hosted deployment paths, including provisioning, upgrades, migrations, configuration changes, and production escalations. Build automation for Terraform runs, customer environment updates, upgrade workflows, secret rotations, and migration steps. Improve observability by turning health signals into clear status, likely causes, and recommended actions. Design safer deployment, upgrade, and rollback paths for Cloud and managed customers. Partner with product engineers on infrastructure requirements for new Retool products and write docs, runbooks, and migration guides. Lead through ambiguity, make careful risk calls, and communicate clearly while priorities change quickly.
Required Qualifications
- Deep experience operating production infrastructure in AWS
- Experience improving reliability for customer-facing SaaS systems
- Strong Kubernetes fundamentals
- Real Terraform or infrastructure-as-code experience
- Good operational judgment around databases, especially Postgres
- Experience building or operating observability systems
- Programming ability in a language such as Go, Python, TypeScript, Java, or Ruby
- Clear written communication
- Comfort working directly with customer-facing teams and, when useful, customers themselves
Desired Qualifications
- You will do well here if you like infrastructure that sits close to real customer pain
- We value SREs who are ambitious, curious, energetic, and careful with the details
- The work needs SREs who can get their hands dirty, tell the truth about tradeoffs, and leave the system better than they found it
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.