Site Reliability Engineer
RemoteCosta Rica
Job Summary
Conduct service resiliency, performance tuning, and system design across Backcountry's multi-cloud platform while owning end-to-end responsibility for critical incidents. Drive resolution of outages through methodical postmortems and leverage AI-assisted engineering tools to automate toil and ship fixes across infrastructure and application repositories. Monitor system health and capacity proactively, building observability instrumentation and SLI/SLOs for services. Partner with developers and architects to implement best practices, participate in FinOps initiatives for capacity planning, and rotate through on-call support.
Required Qualifications
- 3+ years of experience supporting containerized production services, preferably running Kubernetes
- 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.)
- 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus)
- Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable
- Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end
- Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation
- Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js)
- Experience managing Linux (any major distribution) in production environments
- Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.)
- Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work
- Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation
- Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize)
- Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice
- Bachelor's degree in computer science or similar, or equivalent experience
- Advanced-level English communication skills, both verbal and written
Desired Qualifications
- Azure/Entra exposure
- Experience using AI coding assistants (Claude Code, Codex, GitHub Copilot) to build fixes, write automation, and ship application and infrastructure code improvements
- Familiarity with PCI-scoped or other regulated environments
- Previous experience working in ecommerce environments
- Professional certifications: GCP, CKA, or AWS
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.