DevOps Engineer & AI Model Evaluator for Frontier Coding
Summary
Evaluates frontier AI coding models by running infrastructure engineering assessments on cloud, Kubernetes, CI/CD, and observability, with sprint-based task compensation.
Mercor partners with a leading AI research lab to support a Frontier Code Agents project. Contributors evaluate frontier AI coding models through structured technical assessments focused on infrastructure engineering workflows and model evaluation.
Role emphasizes reviewing cloud platforms, Kubernetes, CI/CD, observability, and automation while applying engineering judgment to realistic scenarios. Sprint-based work runs 12–24 hour stretches with compensation per accepted task.