Machine Learning Engineer (Speech/Audio)
Summary
Build and fine-tune speech and audio models (ASR/SpeechLLM) using PyTorch and large-scale data pipelines to improve accuracy for multilingual and domain-specific speech.
Machine Learning Engineer (Speech/Audio) at plaud
About the role Plaud creates hardware and software interfaces that transform human conversation into structured intelligence. We are seeking an engineer to join our Global Product R&D Center to refine our audio processing and speech recognition capabilities.
Key facts
- Location: Singapore
- Engagement: FullTime
- Team: Global Product R&D Center
What you'll do
- Manage large-scale audio and speech data pipelines, including collection, cleaning, filtering, labeling, and augmentation.
- Collaborate with senior engineers to develop hotword mining strategies based on ASR outputs.
- Fine-tune and evaluate SpeechLLM and general LLM models to improve accuracy for code-switching, proper nouns, and industry-specific terminology.
- Adapt models for specific domains using targeted datasets.
- Create evaluation frameworks and test sets to benchmark internal models against commercial and open-source alternatives.
Requirements
- Minimum 1 year of professional experience in machine learning, speech technology, or large-scale data engineering.
- Proficiency in Python and PyTorch.
- Experience with distributed data processing tools such as Ray or Spark.
- Demonstrated background in at least one of the following: ASR or SpeechLLM training/fine-tuning, general LLM/ML model training, or managing large-scale multimodal data pipelines (TB-scale or tens of thousands of hours of audio).
Nice to have
- Experience with SpeechLLM or speech SSL concepts, or familiarity with models like Qwen3-Omni or StepAudio.
- Background in contextual biasing, hotword development, or code-switching ASR improvements.
- Authorship of patents or publications in top venues like ICASSP or Interspeech.
- Previous experience managing data workstreams for projects involving hundreds of thousands of hours of speech.
Skills & Tools
- Python
- PyTorch
- Spark
- Ray
- ASR/SpeechLLM
- Large-scale data pipelines
Practical notes
- Compensation includes an Employee Stock Ownership Plan (ESOP).
- Benefits include medical insurance and WICA coverage.
- Employees receive top-spec laptops, high-performance workstations, and Plaud hardware.
- The company provides access to frontier AI tools including Claude Code and Gemini.
- This is an on-site role based in Singapore.