Principal Applied Scientist, Agentic AI
$191,300–$305,700 year
RemoteUnited States
Job Summary
Define and drive the end-to-end applied science strategy for advanced reasoning and long-running agent systems, balancing innovation, customer value, reliability, safety, and business priorities. Advance agent capabilities through agent architecture and harness design, context and memory engineering, test time scaling, post-training, evaluation, and data flywheel development. Establish rigorous evaluation frameworks for agentic systems, including task- and trajectory-level evaluation, long-horizon benchmarks, controlled experiments, ablation studies, and online measurement. Partner with applied scientists, machine learning engineers, software engineers, product leaders, and senior stakeholders to define the science roadmap for Zillow's agentic experiences. Influence executive stakeholders by communicating scientific insights, technical trade-offs, risks, and recommendations that shape product and platform strategy. Stay at the forefront of research in agentic AI and foundation models, selectively introducing advances that materially improve customer experiences. Represent Zillow's work through publications, conference presentations, open-source contributions, patents, and external technical engagement.
Required Qualifications
- PhD degree in Computer Science, Machine Learning, Statistics, Mathematics, or a related field, or equivalent experience demonstrating comparable depth of technical expertise, scientific leadership, and impact in production AI systems
- 8+ years of industry experience building and deploying large-scale, high-impact AI systems
- deep expertise in large language models and agentic AI, with experience turning research ideas into reliable customer-facing systems
- Experience developing AI agents with capabilities such as multi-step reasoning, long-running execution, and proactive behavior
- Extensive hands-on experience across agent architecture and harness design, context and memory engineering, test-time scaling, and post-training
- Extensive experience establishing rigorous evaluation frameworks for complex AI systems, including benchmark design, experimentation, measurement of reliability and end-to-end impact
- proven track record of leading complex, ambiguous, cross-functional initiatives from scientific framing and prototyping through production deployment, monitoring, and measurable customer or business impact
- Demonstrated technical leadership, mentorship, and ability to influence across science, engineering, product, and executive audiences
- science leadership through high-impact production systems, patents, publications, open-source contributions, conference presentations, or equivalent industry impact
- Must be able to lift 50 lbs
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.