Freelance AI Evaluation Engineer — Developer Task Architect
Summary
Designs and evaluates AI coding agents by creating developer tasks, prompts, and tests in Python/JS to ensure robust, fair, and solvable evaluations.
Mindrift is building a dataset to evaluate AI coding agents by creating realistic developer environments, tasks, and tests. You will craft prompts, define success metrics, and ensure the evaluation is solvable by an AI agent.
You will write tests that accept all valid solutions and iterate based on QA feedback to make the process fair and robust. This role requires 5+ years of software development and strong Python/JS stacks, with English proficiency at least B2+.