Sur La Table logo
Sur La TablePosted 1 month ago

Site Reliability Engineer

RemoteCosta Rica

Full TimeLarge

Job Summary

Conduct service resiliency, performance tuning, and system design across Backcountry's multi-cloud platform while owning end-to-end responsibility for critical incidents. Drive resolution of outages through methodical postmortems and leverage AI-assisted engineering tools to automate toil and ship fixes across infrastructure and application repositories. Monitor system health and capacity proactively, building observability instrumentation and SLI/SLOs for services. Partner with developers and architects to implement best practices, participate in FinOps initiatives for capacity planning, and rotate through on-call support.

Required Qualifications

  • 3+ years of experience supporting containerized production services, preferably running Kubernetes
  • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
  • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
  • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
  • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
  • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
  • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
  • Experience managing Linux (any major distribution) in production environments
  • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
  • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
  • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
  • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
  • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
  • Bachelor's degree in computer science or similar, or equivalent experience
  • Advanced-level English communication skills, both verbal and written

Desired Qualifications

  • Azure/Entra exposure
  • Experience using AI coding assistants (Claude Code, Codex, GitHub Copilot) to build fixes, write automation, and ship application and infrastructure code improvements
  • Familiarity with PCI-scoped or other regulated environments
  • Previous experience working in ecommerce environments
  • Professional certifications: GCP, CKA, or AWS

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce