Senior Staff Research Scientist | Voice
On-siteLondon, England, United Kingdom
Job Summary
Lead hands-on research and development across ASR, MT, TTS, and speech-to-speech translation for real-time voice products. Design, train, and optimize large-scale ASR models for multilingual accuracy and ultra-low-latency streaming. Improve cascaded translation pipelines end-to-end, including segmentation and streaming MT inference. Develop real-time TTS models with natural prosody and fast inference. Build end-to-end and LLM-based speech-to-speech translation systems, owning the full lifecycle from prototyping to production deployment. Work closely with engineering teams to integrate models into real-time systems, driving improvements in inference efficiency and robustness to real-world acoustic conditions. Mentor researchers and engineers while establishing strong practices for evaluation and continuous model improvement in production.
Required Qualifications
- Deep expertise in speech, audio, or multilingual ML, particularly in ASR, MT, TTS, end-to-end ST, or large speech models
- A hands-on builder who enjoys training models, running experiments, debugging pipelines, and integrating ML systems into production
- Strong understanding of real-time streaming constraints and how to design models that operate reliably at low latency
- Experience shipping ML models to production, maintaining them at scale, and working with engineers on deployment, monitoring, and serving
- Ability to lead complex research efforts while staying grounded in product impact, user experience, and real-world performance
- Strong coding and experimentation skills (Python, PyTorch/JAX, audio processing libraries)
- Ability to communicate clearly, collaborate across teams, and align research work with product and engineering priorities
- Proven experience mentoring others and elevating technical quality across a fast-moving, applied research team
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.