Staff Applied AI Scientist
Summary
Lead AI quality and optimization for production LLM systems, focusing on preference learning, fine-tuning, evals, and continuous improvement in a high-scale messaging platform.
About the Role
We are looking for a Staff Applied AI Scientist to lead the quality, performance, and optimization of our production AI systems.
This role focuses on improving the core behavior of our AI systems, including prompting, benchmarking, evals, model optimization, fine‑tuning, and distillation. It is not primarily an application engineering or tooling role.
Our products process more than 4 billion messages per year across real customer communication workflows. We have already built the foundation: a multi‑phase training pipeline using production conversation data as the preference signal, including a trained Reward Model and an LLM‑as‑judge evaluation benchmark across hundreds of topics drawn from real production data. You will inherit this work, advance it into the next phase of preference learning, and build the broader system for continuous AI quality improvement around it.
You will define how we measure quality, identify failure modes, improve accuracy and reliability, and make better model decisions over time. You will work closely with product and engineering to turn real customer scenarios into a disciplined system for evaluation and continuous improvement, with a strong focus on quality, latency, and cost at scale.
Drive the next phase of our preference learning and fine‑tuning pipeline
Lead strategy for AI quality, evals, and benchmarking across production
Define and improve key metrics for production AI performance, including accuracy, instruction following, tool‑use reliability, latency, multilingual performance, and cost efficiency
Build repeatable evaluation and feedback loops to improve quality over time
Drive prompt optimization, failure analysis, and model selection based on real‑world performance
Define the roadmap for distillation and longer‑term model optimization
Partner closely with product and engineering to improve production outcomes for customers
Deep experience in LLMs, NLP, deep learning, or applied AI
Strong track record in one or more of: evals and benchmarking, prompt engineering, fine‑tuning, distillation, conversational AI or agent systems, model optimization in production
Hands‑on experience with preference learning, RLHF, or DPO pipelines is a strong plus
Strong technical judgment on how to improve AI systems beyond application‑layer integration
Experience working on production AI systems where quality, latency, and cost all matter
Comfortable operating without a large team: you will do the work, not just direct it
Strong ownership, product sense, and cross‑functional communication skills
Wati operates at large scale in real customer communication workflows, processing more than 4 billion messages per year. We have already invested in the foundational ML infrastructure.
You will not be starting from scratch.
You will be joining at the inflection point where the pipeline is built and the real work of continuous improvement begins.
This is a rare IC opportunity to own AI quality end‑to‑end at production scale, with real data, real customer impact, and a direct line to the product and founding team.