Lead Machine Learning Engineer
Summary
ASAPP is hiring a Lead Machine Learning Engineer (hybrid, Mountain View, 10-12 office days/month) to own and grow the evaluation platform that measures quality, safety, and performance of its agentic AI/LLM systems. Day to day: designing eval methodologies, building annotation/data infrastructure, and mentoring engineers with Python, AWS, and Kubernetes.
The AI Engineering team is responsible for working closely with the research and modeling teams to create state-of-the-art NLP models for specific tasks, and deploy them in a production setting designed to serve our customers at scale. We are looking for a Machine Learning Engineer to help build and evaluate the core intelligence behind our agentic AI systems. This role will play a key part in designing and owning evaluation frameworks that ensure quality, safety, and performance across complex agentic systems.
We're looking for a Lead Machine Learning Engineer to own and grow the evaluation platform that measures quality, safety, and performance across ASAPP's agentic AI systems- the infrastructure that tells us, with confidence, whether a model or agent change is actually an improvement before it reaches customers.
This a hybrid role with 10-12 days of in-office presence per month to balance flexibility with collaboration.
What you'll do
-
Help develop the technical roadmap and architecture for the evaluation platform, from offline benchmarking to online/production monitoring of agentic and LLM-based systems.
-
Design eval methodologies appropriate to different stages of the pipeline: golden/regression test sets, human-in-the-loop review workflows, LLM-as-judge approaches, and automated metrics for task success, safety, and hallucinations.
-
Build the data infrastructure evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and tooling that lets researchers and product teams run and interpret experiments without needing platform team help.
-
Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions
-
Represent the eval platform to stakeholders outside the immediate team- set expectations on what "good" looks like for a model/agent release, and report on platform health and coverage.
-
Stay current with advancements in ML, NLP, voice, and LLM systems, and contribute actively to technical discussions across teams.
-
Mentor and support other engineers through design reviews, feedback, and knowledge sharing.
What you'll need
-
Deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems- not just consuming existing eval tools.
-
Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).
-
Strong architectural skills, with proven experience designing complex, data-intensive software systems and production experience with Python, AWS, Kubernetes, and/or Docker.
-
Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.
-
A Bachelor’s Degree in CS or other related fields
-
Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment for scalability and extensibility.
-
Desire to learn, teach, and collaborate closely with cross-functional peers.
What we'd like to see
-
Experience building and evaluating agentic systems at scale.
-
Experience with voice/audio quality evaluations.
-
Production experience with LLM-centric services (e.g., inference, orchestration, evaluation, monitoring)
-
Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.
-
Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).
-
Knowledge of techniques for optimizing model architectures for faster inference.
-
Experience with AWS, CI/CD, Kafka, Athena
Skills
As published by lever
Which location are you applying for?, Resume/CV, Full name, Email, Phone, Current location, Current company, LinkedIn URL, GitHub URL, Portfolio URL, Other website