Senior Manager, Data Engineer, Clinical Operations
$153,770–$186,327 year
HybridPrinceton, New Jersey, United States
Job Summary
Support cross-functional teams for AI and Data initiatives across Clinical Trial Operations product lines, developing solutions that accelerate data usage in clinical R&D while ensuring robustness and scalability. Collaborate with BI&T partners, clinical study teams, and domain experts to deliver data product development, standardization, and access governance across clinical trial assets. Optimize the data platform for performance and cost-effectiveness using cloud-native parallel processing, Databricks Delta Lake, and Unity Catalog to enforce data governance and manage metadata. Build and deploy GenAI-powered applications leveraging RAG, fine-tuning, and vector embeddings to drive efficiency gains and automate compliance within clinical operations workflows. Provide mentorship to junior analysts and vendors on technical best practices while maintaining comprehensive documentation of processes and knowledge.
Required Qualifications
- 7+ years of cross-domain experience in Data Engineering, Analytics, and AI/ML
- proven hands-on experience implementing and operating data capabilities and solutions in a cloud environment
- Strong stakeholder engagement and communication skills
- ability to influence and drive adoption of data solutions across functional teams
- Hands-on expertise with Databricks, including Delta Lake, Unity Catalog, Databricks Workflows, Mosaic AI, and MLflow
- Databricks certification (e.g., Databricks Certified Data Engineer Associate/Professional)
- Demonstrated expertise in cloud-native data platforms, ETL/ELT pipeline design, data modeling, and semantic analytics for large-scale, complex datasets
- hands-on DevOps experience
- Strong proficiency in Python, SQL, Spark (including PySpark on Databricks), and GenAI frameworks
- hands-on experience with LLM architectures, RAG, prompt engineering, and agentic frameworks
- Proficiency in creating and maintaining optimal data pipeline architecture for large, complex datasets
- semantic modeling within a domain — preferably life sciences or clinical trial operations
- Demonstrated experience delivering production-grade GenAI applications, predictive models, and self-service analytics tools supporting critical clinical business functions
- Working knowledge of LLM and GenAI-driven approaches, including RAG, Chain-of-Thought, fine-tuning, vectorization, agentic frameworks, and prompt engineering techniques for improving the accuracy of LLM-based responses
- functional knowledge of Life Sciences R&D and clinical trial operations
- Familiarity with clinical trial management systems (e.g., Veeva Vault, Medidata Rave, IRT) and their underlying data structures
Desired Qualifications
- Domain knowledge across Clinical Trail Delivery, including feasibility, site selection, forecasting and trial execution is a strong plus
- Commitment to engineering excellence, documentation, and process improvement
- experience with Generative AI (GenAI), Databricks, and semantic technologies to drive innovation and efficiency within our data platforms
- experience with cloud-native parallel processing, Databricks Delta Lake, caching, and partitioning
- experience with Databricks Unity Catalog to enforce data governance, manage metadata, and ensure end-to-end data lineage across clinical trial data products
- experience with self-service data discovery solutions
- experience with RAG, fine-tuning, and vector embeddings to deliver high-quality data discovery and consumption features for clinical operations workflows
- Stay current on technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
- experience providing guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.