Data Science Quality Architect for AI Evaluation
Summary
This role involves evaluating AI system performance on real-world data science tasks by designing grading criteria and providing detailed, reasoned justifications for scoring. The position requires strong independent judgment and the ability to communicate findings clearly to executive stakeholders.
Mercor is partnering with a leading AI research organization to engage experienced data scientists for a project focused on evaluating how well AI systems perform real-world data science work. You will define what excellent work looks like: designing task-specific grading criteria and scoring completed work samples with rigorous, well-reasoned written justifications.
This role emphasizes independent judgment, reproducibility, and clear communication to executive stakeholders, with emphasis on