Senior AI Data Scientist I
$143,500–$203,000 year
On-siteAlameda, California, United States
Job Summary
Build, train, and validate machine-learning models on clinical datasets to transform complex data into analysis-ready deliverables for drug-development decisions. Execute data cleaning, transformation, and standardization across EDC, vendor, and real-world sources aligned with CDISC (SDTM/ADaM) standards. Develop LLM-based workflows for automated TLF review and create interactive dashboards for study-health monitoring. Maintain data pipelines on Databricks and AWS infrastructure while ensuring GxP-compliant outputs and documentation. Collaborate with Statistical Programming, Clinical Data Management, and Clinical Operations to accelerate data-driven insights and support study-level and portfolio-level decisions. Contribute to manuscripts and conference presentations to drive scientific visibility.
Required Qualifications
- Bachelor's degree in Data Science, Computer Science, Statistics, Biostatistics, Bioinformatics, or a related quantitative field and a minimum of 7 years of experience
- Master's degree in Data Science, Computer Science, Statistics, Biostatistics, Bioinformatics, or a related quantitative field and a minimum of 5 years of experience
- Equivalent combination of education and experience
- PhD: No prior experience applying AI/ML methods to structured or unstructured data
- With Master's degree: A minimum of one (1) year of experience applying AI/ML methods to structured or unstructured data
- With Bachelor's degree: A minimum of three (3) years of experience applying AI/ML methods to structured or unstructured data
- Without degree: A minimum of seven (7) years of relevant professional experience, including demonstrated application of AI/ML methods to structured or unstructured data
- Intermediate proficiency in Python (Pandas, NumPy, scikit-learn) for data manipulation and model prototyping
- Intermediate proficiency in R for statistical analysis and visualization
- Basic proficiency in SQL for data querying and transformation
- Intermediate understanding of supervised and unsupervised learning fundamentals, including model evaluation
- Basic familiarity with NLP, text mining and/or time series analysis techniques
- Basic familiarity with LLM APIs and prompt engineering concepts
- Basic knowledge of Databricks notebooks and Delta Lake concepts
- Basic familiarity with AWS cloud services (S3, Lambda, Glue)
- Basic understanding of data pipeline concepts and data integration fundamentals
- Intermediate proficiency with version control (Git/GitHub) and project tracking tools (Jira)
- Intermediate proficiency with BI platforms including Spotfire, Tableau and/or Power BI
- Basic understanding of the clinical development process and regulatory requirements (ICH, GxP)
- Basic familiarity with CDISC data standards (SDTM, ADaM) concepts
- Ability to communicate technical concepts clearly to diverse audiences
- Strong collaboration and teamwork skills in a cross-functional environment
- Attention to detail and organizational skills
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.