AfterQuery — Research Scientist, Post-Training
Summary
Research Scientist designs and runs controlled experiments to measure how AfterQuery’s training datasets impact LLM performance across reasoning, tool use, and domain-specific tasks, then shares findings with partner labs.
AfterQuery — Research Scientist, Post-Training
Type: Full-time | On-site | San Francisco, CA Compensation: $150K–$250K base + profit sharing (total cash ~$250K–$450K) + equity Hiring count: 3 (client hiring 2–3 Research Scientists) Visa sponsorship: None available Reports to: Not specified on role page
About AfterQuery
AfterQuery builds the training data and evaluation infrastructure that frontier AI labs use to improve their models. They partner with leading labs to design high-signal datasets and run rigorous evaluations that go beyond static benchmarks. Post–Series A, they've raised $30M at a $300M valuation, with a founding team out of Jane Street, Citadel, Google, Goldman Sachs, and Stanford AI Lab. Small, early team where individual contributors have direct impact on how the next generation of models learns and improves.
Founded: 2025 | Team size: 11–50 | Funding: $30M raised at a $300M valuation Founding team: Jane Street, Citadel, Google, Goldman Sachs, Stanford AI Lab Industry: AI / ML — training data & evaluation infrastructure Website: afterquery.com Office: San Francisco, CA
Source note: Funding, valuation, and founding-team pedigree come from the role's outreach template. The company card tags the industry as "Consumer Tech" — treated as a likely mis-tag given the role description.
Why Candidates Should Join
- Frontier-lab leverage: Work directly shapes the datasets leading AI labs use to train next-gen models.
- Strong backing & pedigree: $30M raised at a $300M valuation; founding team out of Jane Street, Citadel, Google, Goldman Sachs, and Stanford AI Lab.
- High cash comp: Base plus profit sharing pushes total cash to ~$250K–$450K, with equity on top.
- Build, don't theorize: Experimental, high-leverage IC work at the edge of model development — not a pure-research seat.
Intake Call Summary
- Not provided on the role page.
The Role
Prove that AfterQuery's data works — design and run training experiments that isolate the impact of their datasets on model behavior (SFT and RL post-training), and turn the results into defensible evidence for partner labs.
What You'll Be Doing
- Run controlled SFT and RL experiments to measure the impact of the datasets on model performance.
- Quantify lift across capabilities — reasoning, tool use, long-horizon tasks, and domain-specific workflows.
- Share findings directly with partner labs to deepen relationships and drive sales.
- Collaborate with internal SPLs to iterate on data quality based on results.
- Work closely with the other Research Scientists to build shared experimental infrastructure and benchmarks.
Tech stack: Not specified (LLM post-training — SFT, RL).
Requirements
- Run controlled SFT and RL experiments to measure dataset impact on model performance
- Quantify lift across capabilities including reasoning, tool use, long-horizon tasks, and domain-specific workflows
- Communicate findings with partner labs to drive sales
- Work with internal SPLs to iterate on data quality based on experimental results
- Strong familiarity with LLM training and evaluation methodologies
- Design lightweight experiments and extract actionable insights from messy results
- Work across multiple domains including finance, software engineering, and policy
Green Flags
- Has run controlled post-training experiments end-to-end, can point to a specific data intervention that shifted model behavior in a measurable way
- Comfortable reading messy experimental results, doesn't need clean data to find signal
- Strong quantitative instincts paired with SWE ability, can actually ship the experiment, not just design it
- Has worked adjacent to or inside frontier labs or eval orgs — understands what "high signal data" actually means in practice
Red Flags
- PhD-only researcher profile with no shipping track record, role explicitly prefers pre-PhD builders
- Wants to focus on a single domain — the work spans finance, code, policy, and enterprise workflows
Additional Context From Role Body
Must-Have (body):
- Strong familiarity with LLM training and evaluation methodologies, including SFT and RL post-training
- Genuine obsession with how data structure, selection, and quality drive model behavior
- Ability to design lightweight experiments, move fast, and extract actionable insights from messy results
- Comfort working across domains — finance, software engineering, policy, and more
- Undergrad or master's research background; pre-PhD candidates preferred
Nice-to-Have (body):
- Prior work or internship at an RL environment company, AI safety org, or benchmarking org (METR, Artificial Analysis, or equivalent)
- Experience running controlled training experiments end-to-end
- Published research on model evaluation, post-training, or data curation
- Strong SWE chops alongside research instincts
Role Details
Salary$150K–$250K baseProfit sharingTotal cash ~$250K–$450KEquityCompetitive (unspecified %)On-site policyOn-site, San FranciscoVisa sponsorshipNone availableEmployment typeFull-timeLocationSan Francisco, CA
Screening Questions
- None provided on the role page.
- Contrario submission form requires (Required Candidate Q&A): (1) LinkedIn profile; (2) legally authorized to work in the country of application?; (3) will you now or in the future require visa sponsorship?
Interview Process
Stage 1 — Resume / project screen — Initial review of resume and projects. Stage 2 — Team interview — Interview with the AfterQuery team. Stage 3 — Take-home — Up to $100 in API credits comped. Stage 4 — Take-home review Stage 5 — In-person work trial (2 days) / Onsite Stage 6 — Offer Extended Stage 7 — Candidate Hired — Candidate accepts and starts.
(Contrario pipeline stages: Pending Approval Initial Screen Take Home Take Home Review Onsite Offer Hired.)
Ideal Companies & Backgrounds
- Not provided on the role page. Nice-to-have signals name RL-environment companies, AI safety orgs, and benchmarking orgs (e.g. METR, Artificial Analysis, or equivalent).