Aloha Consulting Group logo
Aloha Consulting GroupPosted 1 week ago

AI Engineer (Voice & Speech)

On-siteHanoi, Hanoi, Vietnam or Ha Noi, Ninh Binh, Socialist Republic of Vietnam

Full TimeSmall

Job Summary

Conduct end-to-end Speech AI research and development, including exploring prototypes, fine-tuning models, and designing rigorous experiments to improve accuracy and latency. Build and optimize production-ready audio AI pipelines for Automatic Speech Recognition, Text-to-Speech, and voice interaction systems, integrating them into cloud, mobile, and edge deployment environments. Collaborate with Product and Engineering teams to develop AI-powered voice assistants and natural conversational experiences, while managing model serving, inference optimization, and continuous monitoring of real-world performance.

Required Qualifications

  • At least 3 years of experience developing AI/ML systems or AI-powered products
  • Strong Python programming skills
  • Hands-on experience with deep learning frameworks such as PyTorch, TensorFlow, or equivalent technologies
  • Proven experience deploying AI models into production environments
  • Solid understanding of speech processing, audio machine learning, and deep learning techniques
  • Hands-on experience with at least one area of Speech AI, including Automatic Speech Recognition, Text-to-Speech, Audio AI, or Voice AI applications
  • Ability to design and implement end-to-end AI and machine learning workflows
  • Experience conducting model experimentation, evaluation, benchmarking, and performance optimization
  • Ability to translate research concepts into practical and scalable engineering solutions
  • Strong analytical and problem-solving skills
  • Product-oriented mindset with the ability to balance model performance, engineering constraints, and user experience
  • Ability to collaborate effectively with cross-functional engineering and product teams

Desired Qualifications

  • Experience working with modern Speech AI foundation models, frameworks, or technologies such as Whisper, wav2vec 2.0, NVIDIA NeMo, XTTS, VITS, SpeechBrain, Kaldi, or equivalent solutions
  • Experience deploying AI models on mobile or edge devices
  • Experience with real-time audio streaming and processing systems
  • Knowledge of speaker recognition and speaker diarization
  • Experience with voice cloning technologies
  • Strong knowledge of digital audio and signal processing
  • Experience integrating Large Language Models into AI applications
  • Experience developing conversational AI or intelligent voice interaction systems

Hiring someone like this?

Get your role in front of qualified candidates on Sorce.

Get started

Apply to this job in one click with Sorce

Apply on Sorce