System Engineer
On-siteChung-ho, Taipei, Taiwan
Chung-ho, Taipei, TaiwanOn-siteFull TimeEnterprise
Full TimeEnterprise
Job Summary
Build and deploy Cluster, Storage, HPC, AI, Data Center, and Cloud infrastructures onsite while troubleshooting hardware and software issues in rack cabinets. Conduct performance testing and benchmarking for servers, GPUs, and HPC environments, then analyze results to identify bottlenecks and optimize system performance for AI/ML workloads. Design and configure high-speed network topologies using Python scripts to automate testing, monitoring, and system optimization. Document complex test procedures and standard operating procedures for servers, networks, and clusters.
Required Qualifications
- Bachelor's or Master degree in Computer Science or equivalent work experience
- 3+ years of proven experience in a HPC/AI or Cloud/Network management
- In-depth knowledge of Cloud/HPC/AI deployment and testing
- Strong problem-solving and decision-making abilities, with a proactive approach to identifying and resolving issues
- Excellent communication skills, both verbal and written, with the ability to collaborate and build strong relationships with stakeholders at all levels
- Programming experience with Python, Ansible and Linux shell scripting
- Programming experience with web applications, including frontend or backend
- Familiar with Intel/AMD/NVIDIA development toolkits like CUDA, oneAPI, ROCm
- Understanding of AI/ML frameworks (e.g., PyTorch, TensorFlow) and deployment requirements for LLMs
- Familiar with the day-to-day operational support for Cluster, Storage, HPC, AI, Data Center and Cloud infrastructures
Desired Qualifications
- It's a plus if you have CCNA/CCNP certificates
- Positive attitude, desire to learn, time management, and strong interpersonal skills
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.