Senior ML engineer on Apple's Siri Speech Evaluation team who owns the datasets, metrics, and automated judges used to evaluate speech LLMs (ASR, TTS, real-time conversational models) for accuracy, robustness, and conversational quality before they ship. Core stack is Python, large-scale data pipelines (e.g., Spark), and LLM/human evaluation methods.
Sign in to see your match