Centific logo
CentificPosted 1 month ago

Sr. Cloud Server Operations Engineer-1

On-siteSingapore, Singapore

Full TimeSenior LevelBachelors DegreeLarge

Job Summary

Lead cloud infrastructure operations and large-scale server fleet management across cloud environments, driving initiatives for scalability, reliability, and automation. Troubleshoot large-scale infrastructure incidents and collaborate with cloud providers to resolve operational challenges. Manage operational data platforms, design automation workflows, and standardize processes while mentoring engineers and leading knowledge-sharing programs. Monitor workloads, handle on-call activities, and review performance metrics to identify bottlenecks. Host operational meetings and support cloud vendor governance. Requires 5+ years of experience in cloud infrastructure, strong scripting skills, and familiarity with Terraform or Ansible. Preferred experience includes GPU infrastructure and RDMA networking.

Required Qualifications

  • Bachelor's Degree in Computer Science, Electrical Engineering, or related fields
  • At least 5 years of experience in cloud infrastructure operations, server operations, or large-scale infrastructure environments
  • Strong leadership, ownership, and decision-making capabilities in high-pressure operational environments
  • Strong communication and cross-functional collaboration skills in English
  • Deep understanding of cloud infrastructure operations, server lifecycle management, and large-scale operational ecosystems
  • Experience working with public cloud platforms such as Oracle, Amazon Web Services, Google, or Microsoft
  • Strong experience with automated provisioning technologies, bare metal lifecycle management, and infrastructure deployment pipelines
  • Strong scripting and automation capabilities using Shell, Python, or infrastructure APIs
  • Familiarity with automation and infrastructure management tools such as Terraform, Ansible, GitLab CI/CD, or cloud SDKs
  • Strong Linux troubleshooting and infrastructure diagnostic capabilities
  • Strong understanding of networking concepts including TCP/IP, subnetting, VLANs, DNS, IPv6, routing, and cloud networking architectures
  • Experience with infrastructure monitoring, incident management, operational governance, and service reliability initiatives
  • Strong documentation, workflow standardization, and operational process management capabilities

Desired Qualifications

  • Mandarin
  • Experience supporting GPU infrastructure, large-scale AI clusters, firmware lifecycle management, or RDMA networking
  • Experience leading cloud operational programs, vendor management, or infrastructure transformation initiatives

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce