Data & AI - Data Engineer
HybridSingapore, Singapore
Job Summary
Build and operate the data foundation for client artificial-intelligence and analytics solutions by designing ingestion pipelines, lakehouses, and semantic layers that transform fragmented data into authoritative, audit-ready sources. Engineer batch and streaming integrations from ERP, clinical, and operational systems into governed platforms, preparing curated datasets, embeddings, and retrieval corpora for AI agents and models. Own end-to-end data quality, lineage, and master data management, embedding privacy controls and compliance with Singapore's Personal Data Protection Act and Asia-Pacific residency rules. Partner with AI engineers and stakeholders to establish reusable pipelines and accelerators that scale from bespoke builds to repeatable delivery.
Required Qualifications
- Approximately 5–10 years of strong, current, hands-on data-engineering experience building and operating production data pipelines and platforms
- Experience delivering data work end-to-end to a production standard, including designing, building and operating governed data pipelines and platforms
- Client-facing or executive-stakeholder delivery experience, including the ability to explain data design and technical trade-offs clearly to both technical colleagues and non-technical stakeholders
- Expert Python and SQL skills
- Production experience with a transformation and modelling framework, particularly dbt, and disciplined, tested, version-controlled data transformation
- Hands-on delivery experience with Snowflake and/or Databricks
- Working understanding of open table and lakehouse storage formats such as Apache Iceberg and Delta Lake
- Experience with dimensional modelling and data-vault modelling for analytics
- Production experience with a data-orchestration tool such as Dagster, Apache Airflow or dbt Cloud
- Experience with batch and streaming ingestion
- Understanding of the distinction between event transport, such as Apache Kafka, and stream processing, such as Apache Flink
- Experience with change-data-capture patterns
- Experience designing semantic layers using tools or approaches such as Cube, the dbt Semantic Layer or LookML to establish consistent, governed business logic across business-intelligence and artificial-intelligence consumers
- Awareness of ontologies and knowledge graphs for structured knowledge representation in artificial-intelligence-enabled systems
- Familiarity with the Resource Description Framework data model, SPARQL, property-graph databases such as Neo4j, or comparable knowledge-architecture technologies
- Experience with data-quality and validation frameworks such as GX Core, formerly Great Expectations, dbt tests, or comparable tools
- Experience with master-data and entity-resolution techniques
- Experience with data cataloguing, lineage and data-contract tooling such as Atlan, Collibra, OpenMetadata, or comparable platforms for production-grade, audit-ready data
- Hands-on delivery experience on at least one major cloud platform — Microsoft Azure, Amazon Web Services or Google Cloud Platform — and its associated data services
- An appreciation of infrastructure-as-code, such as Terraform, and cost-aware platform design
- A working understanding of how data feeds artificial-intelligence systems, including embeddings and vector stores such as pgvector, Pinecone, Weaviate or Qdrant
- Experience or working knowledge of chunking and retrieval-corpus preparation for retrieval-augmented generation
- Understanding of the difference between analytics-grade and artificial-intelligence-grade data preparation
- Experience building authoritative, lineage-traceable data foundations with clear controls, reproducibility and retained evidence
- A practical grasp of data governance, privacy and residency across the Asia-Pacific markets the practice serves, not Singapore alone
- Working awareness of Singapore's Personal Data Protection Act and how its cross-border-transfer requirements affect data architecture and design
- Working awareness of the region's principal data-protection and data-residency regimes and how they differ on cross-border transfer, including Japan's Act on the Protection of Personal Information, South Korea's Personal Information Protection Act, India's Digital Personal Data Protection Act 2023, China's Personal Information Protection Law and Data Security Law, Hong Kong's Personal Data (Privacy) Ordinance, Australia's Privacy Act and Australian Privacy Principles, and the developing South...
- The ability to design pipelines, classification controls and residency controls that hold up across different Asia-Pacific jurisdictions
Desired Qualifications
- Healthcare and life-sciences data experience involving clinical, claims or real-world-evidence data
- Experience with the interoperability standards and clinical ontologies that structure healthcare data
- Familiarity with Fast Healthcare Interoperability Resources (FHIR), itself a Health Level Seven standard, and legacy Health Level Seven version 2 messaging
- Familiarity with healthcare terminologies and ontologies such as Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT), the International Classification of Diseases (ICD-10 and ICD-11), Logical Observation Identifiers Names and Codes (LOINC), RxNorm for medicines, and the Medical Dictionary for Regulatory Activities (MedDRA) for pharmacovigilance
- Experience with the Observational Medical Outcomes Partnership (OMOP) Common Data Model for harmonising real-world evidence across sources
- Familiarity with Digital Imaging and Communications in Medicine (DICOM) for medical imaging
- Office-of-the-Chief-Financial-Officer data experience involving financial close and reporting, consolidation, controls or risk data — CFGI's core buyer context
- Experience building data foundations in a governed, audit-ready environment
- Experience contributing reusable data accelerators to a growing practice
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.