Machine Learning Scientist, Pretraining
$200,000–$330,000 year
HybridEmeryville, California, United States
Job Summary
Design and develop state-of-the-art autoregressive, diffusion, and representation learning models for protein design. Collaborate across machine learning and protein design teams to adapt pretraining techniques from other domains to biomolecular deep learning. Architect, implement, and optimize core infrastructure supporting the pretraining of protein language models. Curate datasets and design evaluation tasks for generative models while implementing, analyzing, and interpreting computational approaches. Present results to colleagues during regular update meetings within a collaborative, interdisciplinary team shaping the company's scientific and strategic vision.
Required Qualifications
- PhD (or equivalent industry experience) in Computer Science, Machine Learning, Natural Language Processing, Applied Math, Computational Biology, Statistics, or a related field
- Experience with conceiving of, implementing, and evaluating novel machine learning and pretraining large scale LLMs or other models in biomolecular domain
- Publications at major machine learning conferences (NeurIPS, ICML, ICLR) or scientific journals (Nature, Science, Nature Biotech, Nature Methods, PNAS) or experience pre-training LLMs at frontier AI/ML labs
- Experience with modern deep learning frameworks such as Pytorch or Jax
- Legal authorization to work in the United States
Desired Qualifications
- Familiarity with foundational biology of proteins and nucleic acids
- Experience developing machine learning models for proteins (language models, structure prediction, design)
- Experience with cloud compute platforms (GCP, AWS, Azure, OCI)
- Previous experience in data extraction and curation from bioinformatics data sources
- Familiarity with wet lab experimental assays and associated limitations
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.