C-BRAIN Data Engineer (Remote) - Neurology
$75,200–$128,800 year
RemoteUnited States
Job Summary
Design and maintain scalable data ingestion pipelines for multi-institutional neurodegeneration datasets, integrating omics, neuroimaging, and longitudinal clinical records into unified analytical frameworks. Develop ETL/ELT workflows using Apache Spark, dbt, and Airflow while implementing automated monitoring, alerting, and robust data quality validation at every stage. Manage cloud infrastructure on Microsoft Azure, optimizing storage and compute resources for AI tool deployment and ensuring compliance with data use agreements and PHI de-identification requirements. Collaborate with the CTO and research teams to harmonize data from sources like NACC and ADNI, maintaining comprehensive lineage documentation and contributing to technical progress reports.
Required Qualifications
- Bachelor's degree in Computer Science, Data Science, Bioinformatics, Engineering, or a closely related field
- Three years of hands-on data engineering experience, including design and development of data pipelines and ETL/ELT workflows in a production or research environment
- Demonstrated proficiency in Python and SQL
- Experience with cloud data platforms: Microsoft Azure preferred (AWS or GCP also acceptable)
- Familiarity with cloud storage, compute, and access control management
- Experience working with complex, multi-source datasets requiring integration, harmonization, and quality validation
- Demonstrated software engineering background: production-grade Python with version control (Git), code review practices, and automated testing
- Engineering discipline and the ability to build maintainable, auditable code
- Experience working with at least two of the following biomedical data modalities: omics (genomics, transcriptomics, proteomics), neuroimaging (PET, MRI), digital pathology, or longitudinal clinical/EHR data
- Demonstrated experience working with neurodegeneration or Alzheimer's disease research datasets
- Familiarity with the neurodegeneration data landscape — NACC, ADNI, and/or AD/ADRD repositories
- Sufficient understanding of the biological context to communicate meaningfully with research scientists
- Biomedical informatics experience without neurodegeneration domain knowledge is insufficient for this role
Desired Qualifications
- Experience working with biomedical, clinical, or research datasets in an academic medical center, research university, or life sciences organization
- Experience with research data repositories such as ADDI, Synapse, Terra, or NACC/ADNI data platforms
- Experience with one or more of: Apache Spark, dbt, Airflow, Azure Data Factory, or equivalent ETL/ELT frameworks
- Familiarity with data governance frameworks, data use agreements, or federated data architectures
- Experience supporting AI/ML or data science teams as a data engineering partner: understanding how data products are consumed by model training and inference pipelines
- Experience with NAIRR or other research cloud computing platforms
- Experience with data catalog tools, data lineage platforms, or metadata management
- Familiarity with de-identification standards and privacy-preserving data techniques relevant to biomedical research
- Master's or PhD in Computer Science, Data Science, Bioinformatics, Biomedical Informatics, or a related field
- Familiarity with agentic AI frameworks and how curated datasets feed retrieval-augmented generation (RAG) or LLM-based co-scientist systems (e.g., LangGraph, DSPy, or equivalent)
- Experience with NLP techniques relevant to biomedical data: named entity recognition, natural language inference, or knowledge graph construction
- Knowledge of graph data structures and graph platforms (Neo4j, Amazon Neptune, or equivalent) for representing multi-modal biomedical relationships
- Track record of cross-disciplinary collaboration between computational and experimental or clinical teams
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.