Bristol Myers Squibb logo
Bristol Myers SquibbPosted 1 month ago

Data Engineer, Clinical Operations

$87,810–$106,399 year

HybridPrinceton, New Jersey, United States

Full TimeEnterprise

Job Summary

Design, build, and maintain scalable, production-grade data pipelines and platform components supporting Cross Study Operations and Specimen Management product lines, including cross-trial data aggregation, specimen tracking, and biobanking workflows. Collaborate with BI&T partners, clinical study teams, and domain experts to support effective adoption of the Data Platform, ensuring robustness, interoperability, and scalability. Develop and operationalize GenAI-powered and NLP-driven applications using Databricks Mosaic AI and MLflow to deliver efficiency gains, specimen traceability improvements, and compliance automation. Leverage Databricks, Delta Lake, and semantic technologies to enforce data governance, manage metadata, and enable self-service data discovery solutions for large, complex datasets. Serve as a go-to technical expert, providing guidance on technical best practices and fostering continuous learning within the data engineering community.

Required Qualifications

  • 5+ years of hands-on experience in Data Engineering, Analytics, and AI/ML
  • Hands-on expertise with Databricks, including Delta Lake, Unity Catalog, Databricks Workflows, Mosaic AI, and MLflow
  • Databricks certification (e.g., Databricks Certified Data Engineer Associate/Professional)
  • Demonstrated expertise in cloud-native data platforms, ETL/ELT pipeline design, data modeling, and semantic analytics for large-scale, complex datasets
  • Hands-on DevOps experience
  • Strong proficiency in Python, SQL, Spark (including PySpark on Databricks), and GenAI frameworks
  • Hands-on experience with LLM architectures, RAG, prompt engineering, and agentic frameworks
  • Proficiency in creating and maintaining optimal data pipeline architecture for large, complex datasets, including semantic modeling within a domain
  • Demonstrated experience delivering production-grade GenAI applications, predictive models, and self-service analytics tools supporting critical clinical or research business functions
  • Working knowledge of LLM and GenAI-driven approaches, including RAG, Chain-of-Thought, fine-tuning, vectorization, agentic frameworks, and prompt engineering techniques
  • Strong stakeholder engagement and communication skills
  • Ability to clearly articulate technical concepts to non-technical audiences
  • Commitment to engineering excellence, documentation, and process improvement
  • Functional knowledge of Life Sciences R&D and clinical trial operations

Desired Qualifications

  • Experience in a life sciences, clinical trial operations, or specimen/biobanking workflows
  • Experience with Generative AI (GenAI), Databricks, and semantic technologies
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with cloud orchestration and containerization
  • Experience with Databricks Delta Lake, caching, and partitioning
  • Experience with Databricks Unity Catalog
  • Experience with semantic modeling
  • Experience with interoperability standards for large, complex cross-study and specimen datasets
  • Experience with data governance, metadata management, and end-to-end data lineage
  • Experience with self-service data discovery solutions
  • Experience with data product development, standardization, testing, and access governance
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with data assets
  • Experience with junior analysts, interns, and vendor resources
  • Experience with technical best practices
  • Experience with business alignment
  • Experience with continuous learning, engineering excellence, and knowledge sharing
  • Experience with cross-functional teams
  • Experience with BI&T partners
  • Experience with business analysts
  • Experience with data engineers
  • Experience with Clinical Operations specialists
  • Experience with Specimen Management professionals
  • Experience with Cross-Study Operations leads
  • Experience with domain experts
  • Experience with data-driven initiatives
  • Experience with the BMS data ecosystem
  • Experience with data product lifecycle
  • Experience with ingesting, storing, processing, governing, and interacting with data
  • Experience with high-quality, scalable data solutions
  • Experience with innovative, reliable, secure, and easy-to-use data ecosystems
  • Experience with cross-study operations and biospecimen workflows
  • Experience with clinical data challenges
  • Experience with cloud platforms
  • Experience with Generative AI
  • Experience with data engineering
  • Experience with cross-study operations
  • Experience with biospecimen workflows
  • Experience with data product teams
  • Experience with data operations
  • Experience with effective adoption of our Data Platform
  • Experience with scalable, production-grade data pipelines and platform components
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with robustness, interoperability with consumer applications, and scalability
  • Experience with cloud-native parallel processing, Databricks Delta Lake, caching, and partitioning
  • Experience with ETL/ELT pipelines and data models
  • Experience with Databricks, Delta Lake, cloud-native tools, semantic modeling, and interoperability standards
  • Experience with large, complex cross-study and specimen datasets in life sciences
  • Experience with hands-on technical solutions for data product development, standardization, testing, lineage, meeting latency requirements, and ensuring access governance
  • Experience with cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness
  • Experience with guidance and mentorship to junior analysts, interns, and vendor resources on Databricks, technical best practices, and business alignment
  • Experience with a culture of continuous learning, engineering excellence, and knowledge sharing across the data engineering community
  • Experience with cross-study operations and specimen management
  • Experience with cross-trial data aggregation, specimen tracking, and biobanking workflows
  • Experience with cross-study clinical R&D programs
  • Experience with consumer applications
  • Experience with latency requirements
  • Experience with access governance across cross-study and specimen management data assets
  • Experience with self-service data discovery solutions
  • Experience with findability, accessibility, and reusability of cross-study operational and specimen data assets
  • Experience with thorough documentation of processes, data structures, and technical solutions
  • Experience with clear technical recommendations and executing solutions effectively across the enterprise
  • Experience with GenAI-powered and NLP-driven applications
  • Experience with efficiency gains, specimen traceability improvements, cross-study insight generation, risk mitigation, and compliance automation
  • Experience with cloud-based GenAI and LLM-powered applications
  • Experience with data product owners, engineers, and data scientists
  • Experience with RAG, fine-tuning, and vector embeddings
  • Experience with high-quality data discovery and consumption features for cross-study and specimen management workflows
  • Experience with Databricks Mosaic AI and MLflow
  • Experience with machine learning and GenAI models at scale within cross-study and specimen management contexts
  • Experience with technology trends in GenAI, RAG, semantic search, Databricks, cloud orchestration, and containerization
  • Experience with emerging best practices to optimize platform performance, scalability, and cost-effectiveness

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce