ShipperHQ logo
ShipperHQPosted 1 month ago

Principal Site Reliability Engineer - Austin, Texas

HybridAustin, Texas, United States

Full TimeSenior LevelSmallE-commerce Software

Job Summary

Own the technical vision and roadmap for ShipperHQ's cloud infrastructure, reliability, and platform engineering initiatives. Design, build, and maintain highly available, scalable, and secure cloud infrastructure in AWS, while architecting Infrastructure as Code standards using Terraform. Lead the design of CI/CD pipelines, observability stacks, and self-service automation to empower engineering teams. Drive infrastructure modernization through containerization and orchestration, partnering with Security to implement compliance controls and governance. Define reliability standards, SLOs, and incident management best practices, then mentor engineers across teams to promote operational excellence. Participate in root cause analysis and continuous improvement efforts for production systems. This role is ideal for someone who enjoys building platforms rather than simply maintaining infrastructure and thrives in a fast-paced, AI-first engineering culture.

Required Qualifications

  • 10+ years of experience in Site Reliability Engineering, Platform Engineering, DevOps, Cloud Infrastructure, or Software Engineering
  • Proven experience designing and operating large-scale, highly available cloud infrastructure in AWS
  • Strong software engineering background with the ability to write production-quality code and automation
  • Expert-level experience with Infrastructure as Code, preferably Terraform
  • Deep experience designing and maintaining modern CI/CD pipelines using GitLab or similar platforms
  • Strong knowledge of Kubernetes, containerized workloads, and cloud-native architectures
  • Extensive experience with observability platforms, distributed tracing, logging, monitoring, and incident response
  • Experience defining and implementing SLOs, SLIs, and reliability engineering best practices
  • Strong understanding of networking, security, Linux systems administration, and cloud architecture
  • Experience supporting high-traffic SaaS applications and mission-critical production environments
  • Excellent problem-solving skills with the ability to simplify complex technical challenges
  • Demonstrated ability to influence technical direction without direct authority while mentoring engineers across multiple teams
  • Experience working in Agile development environments and partnering closely with cross-functional engineering teams

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce