Senior Python Engineer – AI Task & Evaluation Architect
Summary
Designs and evaluates AI coding tasks by creating realistic challenges, defining success criteria, and writing tests to validate diverse solutions.
Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.
You're helping build a dataset to evaluate AI coding agents by creating realistic tasks, defining success criteria, and writing tests that accept diverse valid solutions while rejecting incorrect ones. You set the pace and collaborate with QA to refine tasks.