freehire launches on Product Hunt on 26 August.

Follow →

Senior Machine Learning Engineer - Speech AI

Summary

Build and deploy scalable speech AI systems, training and optimizing ASR, diarization, and analytics models using PyTorch and modern ML toolkits.

About Us

Fano Labs is a fast-growing AI startup building production-grade speech technologies. We focus on automatic speech recognition (ASR), speaker diarization, and speech analytics.

We are seeking a Senior Machine Learning Engineer to build scalable ML systems and help bring advanced speech models from research to production.

Responsibilities

  • Build and maintain data, training, evaluation, and inference pipelines for speech AI.
  • Train, evaluate, and optimize ASR, diarization, and other speech-processing models.
  • Run systematic experiments to improve accuracy, latency, robustness, and scalability.
  • Develop evaluation and error-analysis frameworks for speech systems.
  • Work with research, MLOps, and engineering teams to deploy and maintain reliable ML services.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, Speech/Language Processing, or equivalent experience.
  • Strong Python, PyTorch, UNIX/Linux, Git, and software-engineering skills.
  • Practical experience training, evaluating, and optimizing deep-learning models.
  • Strong understanding of ML fundamentals, algorithms, and data structures.
  • Strong written and spoken English, problem-solving ability, and cross-functional collaboration skills.

Preferred Qualifications

  • Experience in ASR, speaker diarization, TTS, or related audio ML tasks.
  • Experience handling speech/audio data, model evaluation, and performance diagnostics.
  • Familiarity with modern speech-model architectures or toolkits, such as Transformers, Conformers, Whisper, wav2vec 2.0, ESPnet, Kaldi, or SpeechBrain.
  • Experience with ML infrastructure and deployment tools, such as Docker, CI/CD, Kubernetes, Ray, Spark, or model-serving platforms.
  • Publications in speech, audio, or machine learning venues—such as ICASSP, Interspeech, SLT, ASRU are a strong plus.

We Offer

  • Competitive compensation based on experience.
  • Hybrid work flexibility and medical insurance.
  • A collaborative, fast-paced environment working on real-world speech AI products.

See also