Staff Site Reliability Engineer
RemoteCanada
Job Summary
Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform, architecting CI/CD pipelines for rapid, reliable deployments. Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues while driving incident response and post-mortem analysis. Partner with engineering teams to optimize performance, cost, and reliability of backend services, eliminate sources of toil, and lead department-wide compliance initiatives including PCI and SOC standards. Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
Required Qualifications
- 12+ years of software engineering experience
- significant experience in Site Reliability Engineering or DevOps roles
- Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub)
- Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi
- Deep experience with Kubernetes, container orchestration, and service mesh architectures
- Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring)
- Experience designing and managing high-throughput, distributed systems
- Strong problem-solving skills and a growth mindset—comfortable with ambiguity and making high-stakes technical trade-offs
- Excellent communication skills and a demonstrated ability to mentor engineers
- This role is available throughout Canada
Desired Qualifications
- Experience working in a high-velocity, customer-focused environment
- Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript)
- Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space
- Experience implementing security and compliance best practices in the cloud
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.