Middle DevOps/SRE Engineer
RemoteChina or Taiwan
Job Summary
Operate, monitor, and improve production infrastructure across AWS and Azure (incl. China regions), automating deployments with Terraform, Ansible, and CI/CD pipelines. Proactively enhance stability, performance, and security to prevent incidents, while responding to outages within SLA to drive root-cause fixes. Support developers with infrastructure queries and participate in the on-call rotation covering business hours and weekends. The role requires 3+ years of DevOps/SRE experience with strong AWS and Azure skills, including Kubernetes, Docker, and scripting in Bash and Python. Join a globally distributed team maintaining follow-the-sun operations for products trusted by 10,000+ schools and 4 million students.
Required Qualifications
- 3+ years in DevOps / SRE
- hands-on AWS experience
- Strong AWS: EC2, VPC, RDS, S3, ElastiCache, OpenSearch, EKS, ECS, Lambda, ELB (ALB/NLB), CloudFront, CloudWatch, CloudTrail, AWS Config, Secrets Manager, SES, Security Hub, Inspector
- Azure: App Service, Container Apps, Front Door, CDN, Azure SQL Database, Storage Accounts, Virtual Network (VNet), Azure DevOps
- Containers & orchestration: Kubernetes, Docker, Helm, ArgoCD
- Infrastructure as Code: Terraform, Ansible
- CI/CD & version control: GitHub Actions, solid Git workflow
- Scripting: Bash and Python
- Linux administration, networking, firewalls, web servers (Nginx/Apache)
- Monitoring & observability: Prometheus (Kubernetes), Grafana, ELK, Loki, incident management with PagerDuty
- Solid understanding of SDLC and SLA-based incident response
- Strong troubleshooting skills and a structured, methodical approach
- English B2+
- On-call rotation during business hours plus weekend coverage
- Urgent off-hours/night escalations happen only for major incidents
Desired Qualifications
- AWS / Azure certifications
- Alibaba Cloud experience
- Experience with Ruby, Python, C# web application frameworks (deploying and troubleshooting)
- Basic understanding of disaster recovery & database backups (PITR, restore procedures)
- Experience with compliance-driven environments (GDPR, ISO 27001, China data residency)
- Cost optimization / FinOps awareness
- Experience with AI applications
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.