SRE Engineer - Seattle
HybridSeattle, Washington, United States
Seattle, Washington, United StatesHybridFull TimeMedium
Full TimeMedium
Job Summary
Ensure reliability and performance of Plaud.ai's AI products at scale by designing and operating highly available, scalable cloud-native systems for AI workloads. Own production reliability, incident response, and on-call practices while building observability and reliability automation. Define SLOs, SLIs, and error budgets with engineering teams and drive postmortems to improve operational maturity. Partner with product and engineering teams on reliability design to support the next-generation intelligence infrastructure.
Required Qualifications
- 5+ years in SRE, Infra, or Platform Engineering roles
- Strong experience with cloud platforms (AWS/GCP/Azure/OCI)
- Hands-on with Kubernetes and distributed systems
- Experience in on-call rotation and incident management
- Proficient in at least one programming language (Go, Python, Java)
Desired Qualifications
- Experience supporting AI/ML or data-intensive platforms
- Experience of GPU cluster management
- Knowledge of SLO/SLA frameworks
- Experience in fast-growing or global products
- Exposure to multi-region systems
- Strong written and verbal communication
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.