Innodata logo
InnodataPosted 1 month ago

Applied Data Scientist, Health AI Evaluation & Datasets

$150,000–$175,000 year

RemoteInn, Oberösterreich, Republic of Austria

Full TimeDoctorate Or Professional DegreeLarge

Job Summary

Design dataset specifications, taxonomies, and rubrics for training, fine-tuning, and evaluating health-domain models across clinical text, medical images, waveforms, and structured EHR data. Translate customer goals into measurable acceptance criteria while defining sampling strategies, label schemas, and quality thresholds in partnership with clinicians and data scientists. Build statistical and ML checks for bias, representation, and leakage detection, then instrument datasets into evaluation pipelines using rubric-grounded LLM-as-judge prompts and regression suites. Evaluate model behavior for calibration, hallucination, and safety-critical failures, ensuring outputs align with clinical workflow needs and regulatory standards. Own data quality from source intake through delivery, managing PHI/PII handling and compliance documentation. Contribute internal IP including reusable taxonomies, golden datasets, and clinical review playbooks.

Required Qualifications

  • 5+ years of data science experience
  • at least 2+ years with healthcare, clinical, biomedical, payer, provider, pharma, life sciences, or comparable regulated health data
  • Working knowledge of healthcare data and standards: EHR structure, clinical documentation conventions, ICD-10, CPT, SNOMED CT, LOINC, RxNorm, and at least passing familiarity with FHIR, HL7, or equivalent interoperability concepts
  • Hands-on experience designing ML datasets, not just consuming them: writing annotation guidelines, sizing cohorts, setting quality thresholds, designing QA checks, and shipping data that downstream teams can train or evaluate on
  • Familiarity with LLM-based health AI workflows, including prompt design, rubric-based evaluation, retrieval-augmented generation, LLM-as-judge methods, model comparison, and the limitations of automated evaluation in clinical contexts
  • Strong Python and SQL
  • comfort with pandas, scikit-learn, statsmodels or equivalent tools
  • working familiarity with modern LLM tooling such as Hugging Face, evaluation frameworks, prompt development tools, or model APIs
  • Statistical literacy across sampling design, bias and fairness analysis, inter-annotator agreement metrics (Cohen or Fleiss kappa, Krippendorff alpha), confidence intervals, significance testing where appropriate, error analysis, and the ability to push back when a number is being over-interpreted
  • Solid grasp of healthcare privacy, compliance, and governance: HIPAA, de-identification standards (Safe Harbor and Expert Determination), practical mechanics of working with PHI safely, auditability, access control, and documentation fit for high-stakes or regulated AI programs
  • Degree in a relevant field such as biostatistics, epidemiology, computational biology, health informatics, computer science with a health focus, statistics, a clinical degree with quantitative training, or equivalent demonstrated experience

Desired Qualifications

  • Clinical credentials
  • MD, RN, PharmD, MPH, PhD, or health informatics backgrounds

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce