Senior Site Reliability Engineer
$138,700–$138,700 year
HybridBurnaby, British Columbia, Canada
Job Summary
Design, build, and operate scalable multi-cloud and hybrid infrastructure using Terraform, Pulumi, and GitOps workflows across AWS, GCP, and on-premises data centers. Own Kubernetes platforms end-to-end, including cluster lifecycle, networking, and autoscaling, while driving progressive delivery patterns like blue/green and canary deployments. Build and run the full observability stack with Prometheus, Grafana, and Datadog to define SLI/SLO policies and lead chaos engineering exercises. Drive incident response and post-mortems focused on systemic fixes, automate toil through self-service provisioning, and promote SRE practices across studios through reliability reviews and embedded collaboration. This hands-on technical leadership role shapes production infrastructure for global launch windows and live-service events, partnering with network engineers and game studio developers to ensure millions of players remain connected.
Required Qualifications
- 5+ years in SRE, Platform Engineering, or equivalent infrastructure work at production scale
- Deep Kubernetes experience in cloud environments (EKS or GKE preferred) — networking, storage, multi-cluster patterns
- Strong IaC proficiency with Terraform and/or Pulumi; hands-on with Helm, Terragrunt, and GitOps tooling (ArgoCD or GitHub Actions)
- Modern and Legacy Tech: AWS, GCP, VMware, and Bare metal servers
- Server Configuration using Ansible, Puppet, and AWS Systems Manager
- Observability stack experience: Datadog, Prometheus + Grafana, and OpenTelemetry,
- SLI/SLO/error budget fluency — including how to operationalize them inside engineering teams
- Production-quality code in Go, Python, or TypeScript: tools, automation, and internal libraries
- Linux internals, TCP/IP networking, DNS, and TLS — proven enough to debug at the system level
- Incident response and post-mortem leadership with a track record of systemic follow-through
- All candidates must be legally authorized to work in Canada without requiring current or future employer sponsorship
Desired Qualifications
- Live-service game or large-scale consumer internet experience at millions of concurrent users
- Service mesh depth (Istio, Cilium) and advanced Kubernetes networking
- FinOps and managing resources at cloud scale
- Experience with AI and Agentic Development
- Cloud certifications (AWS Solutions Architect, GCP Professional Cloud Architect, CKA/CKS, or equivalent)
- Experience mentoring SREs or leading reliability working groups
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.