TwelveLabs logo
TwelveLabsPosted 3 weeks ago

Senior Machine Learning Engineer, Jockey Core

RemoteSan Francisco, California, United States or South Korea

Full TimeSenior LevelSmall

Job Summary

Lead serving engineering for Jockey Core, from engine selection through production scale-out on Blackwell. Build benchmarks and load tests that replay the agent's real traffic, measuring TTFT and inter-token latency separately. Apply inference optimization techniques including quantization, batching/scheduling, disaggregated prefill/decode, and speculative decoding to hit cost and latency targets. Develop cost models and drive production hardening through autoscaling, capacity planning, observability, and rollback-safe rollouts. Collaborate to ensure model-efficiency gains land as real serving wins and set the technical bar through design reviews. Explore AI-assisted development tools to improve productivity. This role operates within TwelveLabs' Cognition Models team, which sits between perception and the agent system to deliver structured understanding for multimodal AI workloads.

Required Qualifications

  • Significant experience serving and optimizing large-scale LLM inference in production (vLLM, TensorRT-LLM, SGLang, or similar), across techniques like batching/scheduling, quantization, disaggregated prefill/decode, and speculative decoding.
  • Experience designing and operating large-scale distributed systems in high-performance GPU environments.
  • A track record of driving ambiguous technical decisions with measured latency/throughput/cost data.
  • Experience building observability, SLOs, and failure-response for production services, with strong communication skills.

Desired Qualifications

  • Experience contributing to or customizing the internals of an LLM inference server (vLLM, TensorRT-LLM, SGLang, or similar).
  • Understanding of the serving characteristics of compressed (pruned/quantized) models or reasoning/agentic LLMs.
  • Experience designing multi-region/multi-cluster serving infrastructure or large-scale GPU capacity planning.
  • A Master's/PhD in Machine Learning, Computer Science, or a related technical field.

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce