Speech ASR AI Engineer
Summary
Engineer at VinFast's Speech & Language Processing Center building and fine-tuning ASR models (Vietnamese, English, Bahasa) for cloud and on-device (Android/embedded) use in vehicles. Day-to-day: train/optimize speech models, build ML pipelines and serving infra, and optimize inference with tools like ONNX Runtime, TensorRT, and CUDA.
VINFAST is a pioneering electric vehicle (EV) company committed to revolutionizing the automotive industry with sustainable and innovative mobility solutions. As a leading player in the EV market, VinFast is dedicated to delivering high-quality, cutting-edge electric vehicles that redefine the driving experience. Our team consists of passionate professionals driven by a shared vision of creating a greener and more sustainable future through innovation, technology, and excellence.
The Speech & Language Processing (SLP) Center – Automotive AI Development Institute , is responsible for researching, developing, and deploying advanced Agentic AI and VoiceAI solutions, applied in vehicles and extended to other group-wide use cases such as robotics and AI assistants.
We're looking for people with relevant experience, passion, and drive — ready to challenge themselves, keep learning, and thrive under high pressure to help build innovative products.
- Develop, train, and optimize Automatic Speech Recognition (ASR) models for Vietnamese, English, and Bahasa across both cloud and on-device platforms.
- Fine-tune and adapt ASR models for production environments and customer-specific applications, ensuring high accuracy, robustness, and scalability.
- Build, optimize, and integrate AI models into Android and embedded devices, enabling efficient real-time on-device inference.
- Research, prototype, and evaluate state-of-the-art speech technologies, including LLM-based ASR, speech foundation models, and related multimodal approaches.
- Design and build automated machine learning pipelines covering data collection, preprocessing, labeling, training, evaluation, continuous learning, and model release.
- Develop, deploy, and maintain scalable model serving infrastructure for real-time and batch speech AI applications.
- Optimize inference performance by improving latency, throughput, memory usage, and hardware utilization across GPU, CPU, and edge devices.
- Collaborate with product, platform, and engineering teams to deliver production-ready speech AI solutions and continuously improve model quality through data-driven iteration.
- Monitor production model performance, analyze failure cases, and implement strategies for continuous improvement and reliable deployment
Requirements
- Bachelor's degree in computer science, AI, Electrical Engineering, or a related field.
- Strong foundation in Machine Learning / Deep Learning, especially speech and audio processing.
- Experience training and fine-tuning ASR models (e.g., Icefall, WeNet, Kaldi, ESPnet, NeMo).
- Experience deploying AI models and optimizing inference using ONNX Runtime, Triton, TensorRT, or CUDA.
- Experience developing on-device AI models, including quantization, pruning, and mobile inference optimization.
- Proficiency in Python and C/C++.
- Experience integrating AI models into Android applications using Android SDK/NDK and JNI is a plus.
- Familiarity with Linux, Docker, Git, and CI/CD tools (e.g., Jenkins).
- Passion for researching and applying the latest advancements in Speech AI and LLMs.
Benefits
- Competitive salary
- Premium healthcare package, including PVI insurance & annual health check-ups
- 13th-month salary & performance bonuses to reward your contributions
- Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool...
- Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain
To all recruitment agencies: VinFast does not accept agency resumes. Please do not forward resumes to our careers alias or other VinFast employees. VinFast is not responsible for any fees related to unsolicited resumes.
