Semios logo
SemiosPosted 1 month ago

Senior Site Reliability Engineer

$140,000–$160,000 year

HybridBritish Columbia, Canada

Full TimeSenior LevelMedium

Job Summary

Lead delivery of infrastructure projects and perform higher-risk maintenance while contributing to incident resolution and on-call roster participation. Collaborate with product and software development colleagues to improve product resiliency and reliability, and mentor team members across all aspects of SRE work. Use a data-driven approach to identify architectural changes that enhance reliability, performance, and availability, while identifying non-scaling system components and driving targeted solutions. Maintain and improve Service Level Indicators aligned with availability and performance targets, and promote automation to reduce operational overhead. Manage productivity and workload in a work-from-home environment within a high-performing team operating across time zones and global regions.

Required Qualifications

  • 8+ years of experience in DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering roles supporting production cloud environments
  • 3+ years of experience in a senior or technical leadership capacity, with demonstrated ownership of critical production systems and mentoring of engineers
  • Hands-on experience with modern cloud environments (AWS, GCP, or Azure), including deployment, scaling, monitoring, and cost optimization of SaaS applications
  • 5+ years of relevant experience in DevOps, SRE, or infrastructure engineering roles
  • Proven experience implementing and managing observability stacks (e.g., Datadog, Prometheus, New Relic, Splunk) and driving improvements to SLIs/SLOs
  • Experience in incident management, including participation in on-call rotations and leading post-incident reviews with a focus on continuous improvement
  • Have good knowledge of Linux and bash or similar
  • Be versed in the delivery of a SaaS product on AWS, GCP, or Azure
  • Have strong programming skills (Ruby, Python, Go, etc.)
  • Be competent with Terraform or similar Infrastructure as Code (IaC) tools
  • Have experience with Docker, Kubernetes, EKS, or similar technologies
  • Have experience with CI/CD pipelines on Buildkite or similar platforms
  • Be familiar with building delivery pipelines with Buildkite or similar
  • Have strong version control skills with Git
  • Be experienced with increasing monitoring and observability using Datadog or similar tools (New Relic, Splunk, etc.)
  • Demonstrate a strong automation mindset, with a focus on eliminating repetitive tasks through scripting, tooling, and documentation
  • Enjoy delivering quickly and iterating fast
  • Must be able to lift 50 lbs

Desired Qualifications

  • AI/LLM upskilling — comfort using AI and agentic tooling (e.g., Claude Code) to accelerate investigation, automation, and delivery
  • Service mesh (Envoy / Istio) — hands-on experience deploying and operating a service mesh for traffic management, observability, and secure service-to-service communication
  • NATS — experience running or building on NATS (or comparable messaging/streaming systems) for event-driven and distributed architectures
  • AWS (preferred)
  • GCP/Azure
  • Terraform
  • Docker
  • Kubernetes (EKS)
  • Buildkite/GitHub Actions/Jenkins
  • Python/Ruby/Go
  • Datadog/New Relic/Prometheus
  • Git (GitHub/GitLab)
  • strong Linux

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce