Nebius Token Factory logo
Nebius Token FactoryPosted 1 month ago

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

$195,200–$262,200 year

On-sitePalo Alto, California, United States

Full TimeSenior LevelDoctorate Or Professional DegreeLarge

Job Summary

Own focused research projects from hypothesis through production handoff, designing rigorous experiments in efficient LLM and VLM inference with measurable impact. Invent, evaluate, and productionize methods for quantization, distillation, speculative decoding, and model/runtime co-optimization using PyTorch, Triton, or CUDA-adjacent tooling. Partner with MLEs to convert prototypes into usable components, while preparing internal reports, technical blogs, or papers for external credibility. Mentor engineers on experimental design and scientific rigor, collaborating across GPU kernel, backend infrastructure, product, and customer teams to select high-leverage research bets.

Required Qualifications

  • PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field
  • Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas
  • Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly
  • Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs
  • Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis
  • Excellent written and verbal communication
  • Applicants must be authorized to work in the country in which they apply and will be required to provide proof of employment eligibility as a condition of hire

Desired Qualifications

  • First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues
  • Experience deploying ML models or inference optimizations in production
  • Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals
  • Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation when connected to inference quality or efficiency
  • Open-source research artifacts, widely used benchmarks, high-quality technical blogs, or invited talks in efficient AI systems

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce