AI Engineer
Summary
AI Engineer at EV maker VinFast building an in-vehicle conversational AI assistant: fine-tuning and evaluating LLMs (SFT, LoRA, RLHF/DPO), building RAG and multi-agent orchestration, optimizing models for on-device edge inference, and running production MLOps/LLMOps. Core stack: Python, PyTorch, HuggingFace, vLLM, TensorRT/ONNX/TFLite.
- LLM
& Conversational AI: Research, fine-tune (SFT, LoRA/PEFT, RLHF/DPO), and evaluate LLMs
for natural, proactive, and personalized dialogue. Handle multi-intent
understanding and multi-zone control within a single command, and
implement RAG and Knowledge Bases for automotive domains (user manuals,
warranty & maintenance, traffic laws, POIs).
- Multi-agent
Systems: Design
architectures for agent orchestration, function/tool calling, planning,
and memory. Manage routing between domains (vehicle control, navigation,
knowledge, entertainment) and resolve task conflicts.
- Edge AI
& Optimization: Perform
model compression (quantization, pruning, knowledge distillation) and
optimize on-device inference for low latency and offline capability using
TensorRT, ONNX, or TFLite, balancing model quality against hardware
constraints.
- Personalization
& Proactivity: Build
Context Engines and recommendation models to proactively suggest routes,
charging stations, driving modes, HVAC, and entertainment content based on
context and user habits, learning continuously from real-world feedback.
- Safety & Quality: Develop AI Guardrails to control hallucinations,
block sensitive content, and protect personal data. Build evaluation benchmarks
per feature and ensure stable production operation (MLOps/LLMOps).
Requirements
- Experience: Minimum 3 years of hands-on experience
in AI projects, specifically in LLM/NLP and Agentic systems.
- Core Technical Skills: Mastery of Transformer/LLM
architectures, fine-tuning, RAG, function calling, and prompt &
context engineering. Proven experience designing multi-agent
orchestration, tool use, planning, and memory with output-quality control.
- Edge Deployment: Practical experience optimizing and
deploying models on edge/embedded devices using TensorRT, ONNX Runtime, or
TFLite.
- Software Engineering: Proficiency in Python and frameworks
such as PyTorch, HuggingFace, and vLLM. Experience bringing models to
production: MLOps/LLMOps, containerization, model serving, and monitoring.
- Education: Bachelor's degree or higher in
Computer Science, IT, Data Science, Applied Mathematics, or related
fields. Strong technical English proficiency.
- Preferred Skills (Plus): Experience in Speech (ASR, TTS, Voice Cloning,
Wake-word Detection, Voice Biometrics), Computer Vision (object detection,
driver/occupant monitoring, video understanding), or the
Automotive/IVI/Embedded/real-time domain. Experience shipping large-scale LLM/Multi-agent
products, publications at top-tier conferences (NeurIPS, ICML, ACL, CVPR,
INTERSPEECH, etc.), or open-source contributions are highly valued.