Principal Solutions Architect, AI Data Infrastructure
On-siteSingapore, Singapore
Job Summary
Own the reference architecture for the AI data path end to end, covering parallel and software-defined storage, NVMe, memory disaggregation, and KV-cache strategies. Evaluate, benchmark, and select technologies against real AI workloads to determine performance, capacity, and total cost of ownership. Operate as a trusted advisor to customers, sizing solutions, validating requirements, and resolving production issues, or internally shape the platform roadmap for Kubernetes and bare-metal environments. Apply low-latency fabric expertise including RDMA, RoCEv2, and InfiniBand to connect storage, memory, and GPUs while establishing operational best practices for data protection and resilience. Report to the CTO in Singapore on a full-time basis.
Required Qualifications
- 10+ years in storage, memory, HPC, or AI infrastructure roles
- deep, vendor-agnostic expertise across the AI data path
- hands-on experience of one or more leading high-performance storage and memory platforms—systems in the class of WEKA, VAST Data, DDN, or Dell (PowerScale/PowerFlex)
- understanding of the underlying principles well enough that the specific brand is secondary
- fluency in NVMe and NVMe-oF
- fluency in parallel and software-defined storage
- fluency in object and file systems
- knowledge of low-latency network fabrics (RDMA, RoCEv2, InfiniBand) that link data to GPUs
- tracking of next-generation memory directions such as CXL, memory disaggregation, and KV-cache optimisation for LLM workloads
- ability to design and run workload-driven evaluations across bare-metal, virtualised, and containerised environments
- ability to translate numbers into architecture and business cases
- experience integrating storage and memory into GPU clusters (DGX/HGX and similar)
- reasoning about CAPEX, server-count, and power trade-offs
- ability to advise a C-level customer
- ability to author a clear reference architecture or whitepaper
- ability to mentor engineers
- comfort in high-ambiguity environments where the right architecture has to be created rather than copied
- deep command of the full AI data path – storage, memory, and cache
- rigorous, workload-driven benchmarking
- thinking in systems, reasoning fluently across compute, network, and data
- being pragmatic and vendor-agnostic
- optimising for outcomes rather than badges
- operating as a single-threaded owner, driving to the outcome
- communicating with equal ease across engineering teams and executive audiences
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.