AI & Data Engineer, Data Discovery Services
$87,810–$106,399 year
HybridPrinceton, New Jersey, United States
Job Summary
Build and maintain Python pipelines that pull metadata from enterprise data catalogs, enrich it with taxonomy tags, and publish it to the discovery platform. Tune search indexes and construct a semantic knowledge layer using vector embeddings to support retrieval-augmented generation. Investigate data pipeline issues across Databricks and AWS Glue, build API endpoints, and design metadata pipelines that map cross-domain dataset relationships. This role sits within the Data Discovery Services team, where engineering, search, and applied AI collaborate to make enterprise data findable and actionable for teams improving patient outcomes.
Required Qualifications
- Bachelor's degree in computer science, Data Science, Information Science, Engineering, or a related field
- Demonstrated proficiency in data engineering, software engineering, or a related technical discipline, with a track record of delivering production data pipelines
- Proficient Python skills
- Proficiency in Structured Query Language (SQL)
- Experience with Databricks and AWS Glue for data pipelines and transformations
- Solid data engineering fundamentals: extract-transform-load (ETL/ELT) patterns, data modeling, data quality, and pipeline orchestration
- Familiarity with AWS cloud services (S3, Lambda, API Gateway, Glue)
- Experience with OpenSearch or Elasticsearch
- Understanding of metadata management and data cataloging concepts
- Effective problem-solving skills and willingness to learn
- Good communication skills and ability to work collaboratively in a team
Desired Qualifications
- Master's degree preferred
- Experience with semantic knowledge layers, RAG patterns, vector search technologies, and building AI-ready data inventories
- Familiarity with semantic search, embeddings, chunking strategies, relevance tuning, and semantic knowledge metadata (entity relationships, taxonomies, context enrichment)
- Exposure to AI agent patterns, MCP, or large language model orchestration frameworks
- Experience with ontology or taxonomy technologies (such as Turtle, Resource Description Framework, Web Ontology Language, or SPARQL query language) or management platforms
- Familiarity with graph databases or knowledge graph technologies
- Experience with metadata enrichment, data lineage, or data quality frameworks
- Exposure to Azure OpenAI, Google Vertex AI, or Amazon Bedrock
- Experience with Docker, Elastic Container Service (ECS), or CloudFormation
- Prior exposure to pharma or life sciences
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.