AI Engineer (Voice & Speech)
On-siteHanoi, Hanoi, Vietnam or Ha Noi, Ninh Binh, Socialist Republic of Vietnam
Job Summary
Conduct end-to-end Speech AI research and development, including exploring prototypes, fine-tuning models, and designing rigorous experiments to improve accuracy and latency. Build and optimize production-ready audio AI pipelines for Automatic Speech Recognition, Text-to-Speech, and voice interaction systems, integrating them into cloud, mobile, and edge deployment environments. Collaborate with Product and Engineering teams to develop AI-powered voice assistants and natural conversational experiences, while managing model serving, inference optimization, and continuous monitoring of real-world performance.
Required Qualifications
- At least 3 years of experience developing AI/ML systems or AI-powered products
- Strong Python programming skills
- Hands-on experience with deep learning frameworks such as PyTorch, TensorFlow, or equivalent technologies
- Proven experience deploying AI models into production environments
- Solid understanding of speech processing, audio machine learning, and deep learning techniques
- Hands-on experience with at least one area of Speech AI, including Automatic Speech Recognition, Text-to-Speech, Audio AI, or Voice AI applications
- Ability to design and implement end-to-end AI and machine learning workflows
- Experience conducting model experimentation, evaluation, benchmarking, and performance optimization
- Ability to translate research concepts into practical and scalable engineering solutions
- Strong analytical and problem-solving skills
- Product-oriented mindset with the ability to balance model performance, engineering constraints, and user experience
- Ability to collaborate effectively with cross-functional engineering and product teams
Desired Qualifications
- Experience working with modern Speech AI foundation models, frameworks, or technologies such as Whisper, wav2vec 2.0, NVIDIA NeMo, XTTS, VITS, SpeechBrain, Kaldi, or equivalent solutions
- Experience deploying AI models on mobile or edge devices
- Experience with real-time audio streaming and processing systems
- Knowledge of speaker recognition and speaker diarization
- Experience with voice cloning technologies
- Strong knowledge of digital audio and signal processing
- Experience integrating Large Language Models into AI applications
- Experience developing conversational AI or intelligent voice interaction systems
Hiring someone like this?
Get your role in front of qualified candidates on Sorce.