Senior Python Engineer — AI Coding Evaluation Architect
Summary
Design realistic coding tasks and evaluation environments to test AI coding agents' performance on developer workflows.
Mindrift is building a dataset to evaluate AI coding agents—assessing how well a model handles real‑world developer tasks. You will craft challenging tasks within realistic simulated environments: a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history.
The work demands designing tasks from intermediate states, writing tests that accept all valid solutions and reject incorrect ones, and iterating on tasks based on QA