Server Administrator
On-siteSingapore, Singapore
Job Summary
Own end-to-end performance engineering of high-performance server clusters, establishing baselines and tuning compute, virtualisation, and storage layers to eliminate bottlenecks. Lead capacity planning and forecasting for the regional server estate, modelling workload growth and ensuring headroom for expansion without over-provisioning. Run day-to-day operations and maintenance of production server clusters at the hardware and hypervisor layers within a 24/7 operational model, including provisioning, patching, configuration, change, asset, and fault management. Serve as the escalation point for major incidents, providing 24/7 standby on rotation as the last-resort escalation tier, and leading root-cause analysis and driving permanent corrective actions. Build and maintain infrastructure automation and infrastructure-as-code using tools such as Ansible, Terraform, and Python to streamline provisioning and configuration. Design and continually enhance high-availability, clustering, replication, backup, and disaster-recovery architectures to meet mission-critical availability targets.
Required Qualifications
- Diploma/ Degree in computer science, engineering or equivalent experience
- 5+ years of experience in server and compute infrastructure engineering and operations within mission-critical environments (data centre, colocation, cloud or service provider)
- Deep hands-on expertise with HCI and virtualisation platforms (e.g. Proxmox VE with Ceph, VMware vSphere/vSAN, Nutanix, Hyper-V), including cluster design, high availability and live migration, with a vendor-agnostic mindset and familiarity with supporting services such as Proxmox Backup Server and Proxmox Mail Gateway
- Strong understanding of enterprise server hardware, including firmware lifecycle, out-of-band management (iDRAC, iLO, BMC) and facilities considerations such as power, cooling and rack density
- Experienced in administering Linux at the hypervisor and storage layers at scale, with a strong understanding of Ceph (RBD, CephFS, replication and erasure coding) and data protection (backup, replication, disaster recovery)
- Hands-on experience with automation and infrastructure-as-code (e.g. Ansible, Terraform, Python), and with monitoring, observability and capacity-management tooling
- Ability to participate in a 24/7 rotational standby as the last-resort escalation point, and to support planned maintenance windows
- Excellent analytical and problem-solving skills, with the ability to work effectively with clients, senior management, staff and vendors
Desired Qualifications
- Professional certifications, such as VMware VCP/VCAP, Nutanix NCP/NCM, Microsoft (Windows Server / Azure Local) or Red Hat (RHCE) certifications.
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.