AI Solution Architect
On-siteShanghai, Shanghai, China
EXPIREDShanghai, Shanghai, ChinaOn-siteFull TimeSenior LevelEnterprise
Full TimeSenior LevelEnterprise
Job Summary
Design and manage end-to-end AI infrastructure, Kubernetes clusters, and MLOps/LLMOps platforms for enterprise healthcare solutions. Lead the full lifecycle of AI application development, including agentic systems, model serving, and distributed training pipelines. Architect secure, scalable backend systems using domain-driven design, while overseeing hardware selection, network optimization, and cloud-native automation. Drive large-scale project delivery through Agile methodologies, enforce rigorous security and compliance standards, and mentor cross-functional engineering teams.
Required Qualifications
- 10+ years of experience as Solution Architect
- minimum of 2+ years of hands-on experience specifically in AI Architecture or AI Platform design
- Fluent bilingual capabilities in both English and Mandarin
- Deep mastery of the Kubernetes ecosystem and full-lifecycle cluster management
- Extensive experience designing and building MLOps and LLMOps platforms
- Deep knowledge of core ML components including Feature Stores, Model Asset Management, SFT (Supervised Fine-Tuning), and Model Serving
- High proficiency in RAG architecture design
- Hands-on experience developing Agentic applications using mainstream frameworks such as LangChain and LangGraph
- Familiarity with Vector Databases (VectorDB) and Graph Databases (GraphDB)
- Rich experience in High Availability (HA) system design, Disaster Recovery, RBAC, and integrating enterprise-level components (e.g., SSO, LDAP, Bastion hosts)
- Proven ability to ensure that architectural designs strictly adhere to security compliance standards and FinOps governance requirements
- In-depth understanding of Cloud-Native architectures (e.g., AWS, Alibaba Cloud)
- High proficiency in the CNCF mainstream open-source technology stack, with particular expertise in observability tools (e.g., Prometheus, Grafana, OpenTelemetry, ELK)
- A strong platform engineering mindset with extensive hands-on experience using Terraform, Ansible, Packer, and CloudFormation for automated deployment and system security hardening
- Mastery of at least one core backend development language (e.g.Python, Go)
- Expertise in building robust, end-to-end pipelines utilizing tools such as GitLab CI/CD and JFrog
- Strong familiarity with foundational AI hardware (GPU/high-performance CPU) and storage technologies (NAS/SAN)
- Proven ability to conduct performance benchmarking and hardware selection for high-compute clusters
- Profound understanding of data center network architectures, with expertise in planning and configuring out-of-band networks and high-performance RDMA passthrough environments
- Mastery of Linux (RHEL/CentOS) operating systems and high proficiency in configuring the NVIDIA driver full-stack and supercomputing driver toolchains
- High proficiency in Scrum/Agile development processes and necessary tools
- Demonstrated track record of leading large-scale AI platforms or projects and mentoring cross-functional teams to successful delivery
Desired Qualifications
- Experience in the pharmaceutical, life sciences, or top-tier tech industries (with large-scale compute platform exposure)
- Familiarity with GxP relevant architecture design experience
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.