Senior ML Research Scientist, Jockey Core
RemoteSan Francisco, California, United States or South Korea
San Francisco, California, United States or South KoreaRemoteFull TimeSenior LevelDoctorate Or Professional DegreeSmall
Full TimeSenior LevelDoctorate Or Professional DegreeSmall
Job Summary
Drive model compression and efficiency research for Jockey Core, focusing on structured pruning, quantization, distillation, and recovery fine-tuning to build production-efficient reasoning models. Design rigorous evals on the agent's real reasoning and tool-calling behavior by replaying real traffic, and explore post-training techniques like SFT/RL to preserve agentic tool-use. Collaborate closely with serving engineers to translate efficiency gains into real cost and latency wins, while training draft models for speculative decoding where beneficial.
Required Qualifications
- Strong LLM research experience — post-training (SFT/RL), model compression, distillation, or efficient inference
- A track record of independently driving research from ideation to execution, with strong experimental judgment (eval design, rigorous ablations, clear empirical conclusions)
- Strong proficiency in Python and PyTorch
- The ability to communicate and collaborate closely with both researchers and engineers
Desired Qualifications
- Hands-on experience pruning, quantizing, or distilling large models, and an understanding of how compression affects reasoning/agentic behavior
- Experience with large-scale distributed training in high-performance GPU environments
- Experience translating research advances into production ML systems
- A Master's/PhD in Machine Learning, Computer Science, or a related technical field
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.