Principal Software Engineer (Infra/Platform)
$180,000–$235,000 year
On-siteSeattle, Washington, United States
Job Summary
Own the reliability, scalability, and operational health of Gradial's production platform by leading the evolution of Kubernetes, CI/CD, and infrastructure as code. Set standards for designing and shipping reliable systems while building automation tooling that helps engineers move faster. Drive improvements in monitoring, alerting, and incident response, partnering with engineering to identify scaling risks early. Influence the long-term direction of the platform across reliability, security, performance, and cost in a high-growth environment. This hands-on individual contributor leadership role requires 7+ years of experience in platform engineering or SRE with deep expertise in cloud-native architecture and observability.
Required Qualifications
- 7+ years of experience in platform engineering, infrastructure, SRE, DevOps, or related roles with direct ownership of production systems
- Proven success designing and operating production-grade infrastructure in fast-moving, high-growth environments
- Deep expertise in Kubernetes, cloud-native architecture, and container orchestration
- Strong experience with infrastructure as code, GitOps, CI/CD workflows, and modern deployment practices
- Strong command of observability and reliability fundamentals across metrics, logging, tracing, alerting, and incident response
- A track record of leading through influence, making sound technical decisions, and raising the bar across engineering teams
- Learn quickly, actively seek out new challenges, and regularly reconsider 'how it's always been done'
- Want to have real impact with zero bureaucracy
- Have an innate drive for being 1% better than the day before, building towards greatness
- Embrace AI as a core tool for problem-solving, innovation, and scale
- Show customer-obsession (internal or external), high ownership/accountability, and bias for action
- Communicate clearly, directly, with curiosity, and assuming good intentions
- Thrive in fast-paced, hyper-growth environments where building better > maintaining status quo
- AI Literacy as a core competency in hiring decisions (thoughtful application of AI tools in work)
- No over-reliance on AI-generated responses during the interview process
Desired Qualifications
- Familiarity with AI or ML infrastructure, including GPU provisioning, model deployment, or compute-intensive workloads
- Experience supporting cloud or multi-cloud environments with a focus on resilience and scale
- Comfort with TypeScript or Python for internal tooling and operational automation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.