Site Reliability Engineer (SRE)
On-siteLagos, Lagos, Nigeria
Job Summary
Design, implement, and maintain highly available and scalable infrastructure; monitor production systems and proactively identify performance bottlenecks; manage incident response and root cause analysis (RCA); develop automation scripts and tools to improve operational efficiency; implement and maintain CI/CD pipelines; manage cloud infrastructure across AWS and hybrid environments; configure and maintain observability platforms including monitoring, logging, and alerting solutions; define and track SLIs, SLOs, and error budgets; support application deployments and release management; collaborate with Engineering, Security, Data, and Product teams to improve system reliability; perform capacity planning and disaster recovery testing; ensure infrastructure and systems comply with security and regulatory requirements.
Required Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field
- 4–7 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Infrastructure Operations
- Strong knowledge of AWS services (EC2, ECS/EKS, RDS, Lambda, VPC, IAM, CloudWatch)
- Experience with Infrastructure as Code (Terraform, CloudFormation)
- Knowledge of containerization technologies (Docker, Kubernetes)
- Experience with CI/CD tools (GitHub Actions, GitLab CI/CD, Jenkins, Azure DevOps)
- Experience with monitoring tools such as Datadog, Prometheus, Grafana, New Relic, or ELK Stack
- Strong Linux administration skills
- Experience with scripting languages (Python, Bash, PowerShell)
- Understanding of networking, DNS, load balancing, VPNs, and security controls
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.