Encora logo
EncoraPosted 1 week ago

Site Reliability Engineer

RemoteColombia or Costa Rica

Full TimeBachelors DegreeMediumTechnology

Job Summary

Design, build, and operate scalable cloud platforms using Terraform and manage Kubernetes-based containerized environments. Ensure reliability and stability of distributed production systems by defining SLOs, SLIs, error budgets, and alerting standards. Implement observability practices with tools like Datadog and Prometheus while participating in on-call rotations to respond to production incidents. Collaborate with engineering teams to improve automation and drive continuous reliability improvements through postmortem reviews. Requires 6+ years of experience in SRE, Platform, or DevOps engineering with strong skills in Python, Go, or Java. Located in Costa Rica, Peru, Colombia, and Bolivia.

Required Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Systems, Software Engineering, or a related technical field, or equivalent practical experience
  • 6+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, DevOps Engineering, Backend Engineering, or Production Engineering
  • Strong software engineering skills in at least one language such as Python, Go, Java, TypeScript, or C#
  • Strong understanding of distributed systems, microservices, APIs, asynchronous processing, queues, databases, caching, retries, idempotency, and failure modes
  • Experience with cloud infrastructure on AWS, Azure, or GCP
  • Experience with Kubernetes, containers, Terraform or similar IaC tooling, CI/CD pipelines, and Linux-based systems
  • Experience with observability tools such as Datadog, Prometheus, Grafana, OpenTelemetry, CloudWatch, New Relic, Splunk, or Sentry
  • Experience defining and operating SLOs, SLIs, error budgets, alerting standards, dashboards, runbooks, and incident response practices
  • Strong communication skills and experience working across cross-functional teams

Desired Qualifications

  • Cloud, Kubernetes, Infrastructure, Reliability Engineering, Security, or DevOps certifications
  • Experience in logistics, transportation, final-mile delivery, field-service software, routing, dispatch, or fleet operations
  • Experience working with operational SaaS or marketplace platforms
  • Experience driving automation, platform reliability, and operational excellence initiatives

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce