freehire launches on Product Hunt on 26 August.

Follow →

Applied AI Researcher, Agent Systems & Evaluation

Summary

Research and build evaluation pipelines for AI agents, including automated loops, task suites, and reinforcement learning experiments using production data.

- Assess model trustworthiness of evaluators - Build evaluation pipeline for AI agents - Collect eval data from production traces - Construct automated eval loop for every change - Design task suites that reflect real distribution - Design verifier guided selection and escalation strategies - Prevent task suite gaming - Read frontier research and convert into live experiments - Run automated hill climbing for prompt context tools and routing - Run test time scaling experiments - Set noise floors and acceptance statistical standards - Supervise fine tune vision language models - Train with reinforcement learning Perks/Benefits: - Direct access to compute - Hands on research with production data - Work with CEO - Work with engineering leadership

See also