Data Scientist II, ML Infrastructure
$114,297–$235,319 year
RemotePalo Alto, California, United States
Job Summary
Translate research-grade data science workflows into production ML pipelines using Airflow, WandB, and Ray while establishing reusable patterns. Apply causal inference methods like propensity scoring and IPW to address high-stakes measurement questions, building self-serve tooling for non-experts. Partner with ML engineers and product teams to identify opportunities for improved tooling and metrics that unlock step-change improvements in model quality. Leverage metadata and engagement signals to build data-driven frameworks for feature importance and content deindexing. Design centralized ML platform tooling to improve feature and model creation, evaluation, and trust across all models operating at scale.
Required Qualifications
- 2+ years of hands-on experience as an applied scientist, ML engineer, research scientist or software engineer, with significant ML production experience
- Strong Python skills
- experience with PyTorch or equivalent deep learning frameworks
- familiarity with distributed compute (Spark, Ray)
- Deep ML theory knowledge with extremely strong fundamentals that can help us reason about ML models from first principles
- Proficiency in software development best practices including version control, code review, and reproducible ML pipelines
- Experience with workflow management tools (Airflow, Prefect, Jenkins, or similar) for reliable ML pipeline orchestration
- Bachelor's/Master's degree in a relevant field such as Computer Science, or equivalent experience
- This role will need to be in the office for in-person collaboration 3-5 times/quarter
- US based applicants only
Desired Qualifications
- Ray specifically is a strong plus
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.