Associate Director, Principal Data Engineer
$176,500–$246,500 year
On-siteSan Francisco, California, United States
Job Summary
Build AI-enabled research workflows and agents to help scientists search, analyze, summarize, and connect Vir-generated data for TCE research insights, Experimenta QC/QA review, and scientific/regulatory reporting. Design trusted LLM, RAG, knowledge-search, and agentic AI workflows with traceability, validation, governance, and human oversight. Lead development and integration of scientific platforms including SeqAssembler, HPD/dAIsY, SASTRY, and OPAL Miner while building scalable pipelines for bioinformatics, genomics, machine learning, and clinical decision-making. Architect and operate AWS cloud infrastructure for AI/ML, bioinformatics, and large-scale scientific data. Strengthen data quality, metadata, lineage, observability, security, and governance across the platform. Partner cross-functionally to align data, cloud, application, and AI solutions with program needs. Provide senior technical leadership across architecture, engineering practices, mentoring, and roadmap development. Report to the Senior Director, Research, based in San Francisco with three office days per week.
Required Qualifications
- BS/MS in Computer Science, Data Engineering, Bioinformatics, Computational Biology, Engineering, or a related technical
- 12+ years of experience in data engineering, software engineering, scientific computing, cloud architecture, or scientific application
- Strong hands-on experience with Python and/or Java, including scalable data pipelines, APIs, services, and production-grade data
- Experience building AI/ML-enabled scientific applications using leading LLMs such as GPT, Claude, Gemini, or similar models, including RAG, knowledge search, agentic workflows, and governance
- Experience integrating LIMS, ELN, scientific data platforms, or experiment-management systems such as Experimenta to support data capture, analysis, quality review, and scientific
- Strong AWS cloud infrastructure experience, including compute, storage, networking, security, data processing, monitoring, automation, and operational
- Experience with modern cloud-based data platforms and workflow technologies such as Databricks, Snowflake, Redshift, Spark, Airflow, Nextflow, or similar
- Experience supporting bioinformatics, genomics, machine learning, clinical, healthcare, or other scientific data
- Strong knowledge of data modeling, metadata management, data quality, observability, governance, lineage, and platform
- Demonstrated technical leadership, including architecture ownership, roadmap development, mentoring, and cross-functional
- Excellent communication and collaboration skills across Research, Clinical, Data Science, IT, and business
- Must currently be authorized to work for any employer in the U.S.
- Must be available for three days per week in the office
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.