Site Reliability Engineer (SRE)
On-siteHo Chi Minh City, Ho Chi Minh City (HCMC), Vietnam
Job Summary
Conduct production deployments, manage incident management and post-incident analysis, and execute business continuity plans and compliance initiatives. Secure customer data through access management, oversee service migrations from on-premises to multi-Cloud setups, and lead security operations. Participate in follow-the-sun on-call activities with additional compensation for idle time. Architect cloud infrastructure, coordinate migrations, and coach team members to improve SLI/SLO metrics. Join Cloud Operations to protect critical infrastructure for global clients using data-driven decision-making and automation tools.
Required Qualifications
- At least 1 year of relevant working experience
- Linux systems and/or Windows Server
- Containers and orchestration
- Cloud automation (terraform / ansible / others)
- One major Cloud provider (access management, networking, compute, serverless, databases, monitoring)
- Bash, Python as scripting languages, application ecosystem for Node.js / Ruby / Python / PHP / Java
- Software delivery (git, CI/CD)
- Application, Environment and Data security practices
- Monitoring and alerting
- Incident Management and Disaster Recovery
- It is expected that all SREs participate in the on-call activities (follow-the-sun model)
- On-call availability (idle time) is subject to additional compensation
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.