Kong logo
KongPosted 3 weeks ago

Site Reliability Engineer

RemoteItaly

Full TimeMedium

Job Summary

Build and maintain core infrastructure as code using Terraform and Ansible while implementing robust monitoring, logging, and alerting systems to ensure 99.99% uptime. Resolve production incidents through systematic debugging and drive blameless post-mortems to prevent recurrence. Write automation to reduce operational toil and enable self-service, collaborating with developers to embed reliability best practices into the application lifecycle. Contribute to capacity planning, disaster recovery drills, and security hardening processes while participating in a fair on-call rotation.

Required Qualifications

  • Experience operating production workloads on a major cloud provider (AWS, GCP, Azure)
  • Proficiency in at least one programming or scripting language, such as Golang, Python, or Bash
  • Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes)
  • Knowledge of Infrastructure as Code principles and tools (Terraform is a plus)
  • Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins)
  • An understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK)

Desired Qualifications

  • Terraform

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce