Site Reliability Engineer
RemoteItaly
ItalyRemoteFull TimeMedium
Full TimeMedium
Job Summary
Build and maintain core infrastructure as code using Terraform and Ansible while implementing robust monitoring, logging, and alerting systems to ensure 99.99% uptime. Resolve production incidents through systematic debugging and drive blameless post-mortems to prevent recurrence. Write automation to reduce operational toil and enable self-service, collaborating with developers to embed reliability best practices into the application lifecycle. Contribute to capacity planning, disaster recovery drills, and security hardening processes while participating in a fair on-call rotation.
Required Qualifications
- Experience operating production workloads on a major cloud provider (AWS, GCP, Azure)
- Proficiency in at least one programming or scripting language, such as Golang, Python, or Bash
- Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes)
- Knowledge of Infrastructure as Code principles and tools (Terraform is a plus)
- Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins)
- An understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK)
Desired Qualifications
- Terraform
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.