Applied Scientist, AI
$180,000–$260,000 year
HybridSan Francisco, California, United States
Job Summary
Explore messy healthcare data and turn ambiguous product or operational problems into well-scoped machine learning, ranking, optimization, NLP, or LLM-based tasks. Build strong baselines, design honest offline and online evaluations, and run careful error analysis to improve model quality and real-world outcomes. Partner with ML engineering to productionize models reliably and work with clinical stakeholders to validate assumptions and review edge cases. Explain model tradeoffs, uncertainty, and limitations clearly to product, operations, and leadership teams. Pressure-test results for robustness before recommending deployment.
Required Qualifications
- MS or PhD in computer science, statistics, machine learning, applied math, operations research, biomedical informatics, epidemiology, or a related quantitative field
- Exceptional applied experience that substitutes for formal graduate training
- Depth in LLMs, ranking, NLP, uncertainty quantification, causal inference, optimization, or healthcare AI
- Experience working with healthcare data such as claims, EHR, clinical notes, scheduling, utilization, quality, risk, or patient engagement data
- Experience working with PHI, HIPAA-aware systems, or other sensitive regulated data
- Experience collaborating with clinicians, clinical operations teams, or other high-stakes domain experts
- Experience in a startup or fast-moving applied environment where ambiguity, speed, and rigor all mattered
Desired Qualifications
- Shipped models that reached production and had measurable real-world impact
- Knowledge of when traditional ML approaches are likely to outperform LLMs, and when LLMs are the right tool
- Understanding of how ML models work under the hood and can explain them clearly to non-technical stakeholders
- Focus on impact and knowing that the simplest model is often the best one
- Treating evaluation as one of the most important parts of model development
- Noticing when a metric is misleading, incomplete, or disconnected from real-world outcomes
- Catching leakage, bias, and confounding that others miss
- Moving fluidly between modeling, error analysis, stakeholder partnership, and production handoff
- Comfort with ambiguity and ability to adapt modeling approaches to problems that do not come with a playbook
- Balancing scientific rigor with the practical need to ship useful systems
- Experience using AI coding assistants such as Claude Code, Cursor, or similar tools as part of your development workflow
- Built, evaluated, and iterated on machine learning or AI models for real-world use cases
- Turned ambiguous business, product, clinical, or operational problems into measurable modeling tasks
- Designed rigorous offline evaluations, experiments, or analyses that informed production or product decisions
- Worked with messy real-world datasets where labels, outcomes, and causal relationships are imperfect
- Used statistical reasoning, experimental design, and error analysis to understand model performance
- Built models using Python and standard ML or AI tooling such as PyTorch, scikit-learn, NumPy, pandas, Polars, Hugging Face, Matplotlib, or similar
- Compared modeling approaches and made pragmatic decisions about when to use traditional ML, LLMs, heuristics, or simpler baselines
- Communicated model performance, limitations, tradeoffs, and uncertainty to technical and non-technical stakeholders
- Partnered with engineering, product, data, operations, clinical, or domain experts to move models closer to production impact
- Operated with enough engineering depth to run experiments end to end and self-serve deployments or production handoffs when needed
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.