Sr. Cloud Server Operations Engineer-1
On-siteSingapore, Singapore
Job Summary
Lead cloud infrastructure operations and large-scale server fleet management across cloud environments, driving initiatives for scalability, reliability, and automation. Troubleshoot large-scale infrastructure incidents and collaborate with cloud providers to resolve operational challenges. Manage operational data platforms, design automation workflows, and standardize processes while mentoring engineers and leading knowledge-sharing programs. Monitor workloads, handle on-call activities, and review performance metrics to identify bottlenecks. Host operational meetings and support cloud vendor governance. Requires 5+ years of experience in cloud infrastructure, strong scripting skills, and familiarity with Terraform or Ansible. Preferred experience includes GPU infrastructure and RDMA networking.
Required Qualifications
- Bachelor's Degree in Computer Science, Electrical Engineering, or related fields
- At least 5 years of experience in cloud infrastructure operations, server operations, or large-scale infrastructure environments
- Strong leadership, ownership, and decision-making capabilities in high-pressure operational environments
- Strong communication and cross-functional collaboration skills in English
- Deep understanding of cloud infrastructure operations, server lifecycle management, and large-scale operational ecosystems
- Experience working with public cloud platforms such as Oracle, Amazon Web Services, Google, or Microsoft
- Strong experience with automated provisioning technologies, bare metal lifecycle management, and infrastructure deployment pipelines
- Strong scripting and automation capabilities using Shell, Python, or infrastructure APIs
- Familiarity with automation and infrastructure management tools such as Terraform, Ansible, GitLab CI/CD, or cloud SDKs
- Strong Linux troubleshooting and infrastructure diagnostic capabilities
- Strong understanding of networking concepts including TCP/IP, subnetting, VLANs, DNS, IPv6, routing, and cloud networking architectures
- Experience with infrastructure monitoring, incident management, operational governance, and service reliability initiatives
- Strong documentation, workflow standardization, and operational process management capabilities
Desired Qualifications
- Mandarin
- Experience supporting GPU infrastructure, large-scale AI clusters, firmware lifecycle management, or RDMA networking
- Experience leading cloud operational programs, vendor management, or infrastructure transformation initiatives
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.