Site Reliability Engineer
RemoteSpain or Portugal
Job Summary
Design, implement, and maintain cloud infrastructure on AWS and Google Cloud platforms while ensuring the scalability, performance, and reliability of Kubernetes-based distributed database systems. Collaborate with developers to write efficient, production-grade code in Golang for automating infrastructure management and improve CI/CD pipelines, monitoring systems, and deployment processes. Proactively identify system bottlenecks, troubleshoot complex issues across the stack, and develop strategies for disaster recovery, high availability, and fault tolerance. Participate in on-call rotations to support critical production systems and respond to incidents. Implement monitoring, logging, and alerting systems to ensure visibility into system health and performance.
Required Qualifications
- Proven experience as an SRE or DevOps Engineer in a cloud-native environment
- Proficiency with Kubernetes in managing large-scale, distributed systems
- Experience with cloud providers such as AWS and Google Cloud (GCP)
- Solid understanding of networking, security practices, and troubleshooting methods
- Understanding of Linux internals (processes, environment variables etc.)
- Familiarity with containerization technologies (e.g., Docker)
- Knowledge of CI/CD practices and tools (Jenkins, CircleCI, etc.)
- Familiarity with alerting, monitoring and observability tools (e.g., Prometheus, Grafana, ELK stack)
- Strong troubleshooting and problem-solving skills, with the ability to address complex infrastructure issues
- Excellent communication and collaboration skills with a focus on continuous improvement and operational excellence
- Strong ability to self-organize and to work independently as part of a remote team
- Knowledge of version control systems, particularly Git
- Familiarity with programming languages such as Golang or Python
- Location: EU Timezone, preferably within the EU itself (Remote)- Spain, Portugal, etc.
Desired Qualifications
- Experience managing distributed databases or large-scale data storage systems
- Knowledge of security best practices in cloud environments
- Experience with scripting languages like Python or Bash
- Experience with Infrastructure-as-Code (IaC) tools like Terraform is a plus
- Experience working with GitOps
- Strong programming skills in Golang, with experience in developing automation tools, scripts, or services
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.