Agentic Coding Annotator - Online / Offline Tasks
About Turing
Turing is one of the world’s fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Turing helps customers in two ways: working with the world’s leading AI labs to advance frontier model capabilities in thinking, reasoning, coding, agentic behavior, multimodality, multilinguality, STEM, and frontier knowledge; and leveraging that work to build real-world AI systems that solve mission-critical priorities for companies.
Role Overview
We are looking for experienced software engineers to evaluate and improve agentic coding models. You will work with realistic coding tasks, review model trajectories, validate solutions, and provide high-quality technical evaluations. Depending on the assignment, you may also design coding tasks, develop evaluation rubrics, and calibrate tasks for AI model benchmarking.
What does day-to-day look like
- Execute realistic coding tasks within agentic coding environments and evaluate model-generated solutions.
- Review model trajectories, compare outputs, and identify meaningful differences in correctness, reasoning, and behavior.
- Validate solutions by reading code, running tests and commands, checking logs, and inspecting artifacts.
- Write concise, evidence-based rationales for rankings and assessments.
- Follow detailed evaluation instructions, milestones, and quality guidelines consistently.
- For offline assignments, design realistic multi-step coding tasks, calibrate them through testing, and create task-specific evaluation rubrics.
- Identify and escalate broken environments, ambiguous tasks, or evaluation issues with supporting evidence.
Requirements
- 5+ years of professional experience in software engineering, QA, developer tooling, data/ML engineering, or another code-intensive technical role.
- Strong practical proficiency in at least 1–2 programming languages or ecosystems, such as Python, JavaScript/TypeScript, Rust, Java, C/C++, Bash, Haskell, Swift, or SQL.
- Strong ability to read unfamiliar codebases, debug issues, run and interpret tests, reason about edge cases, and determine whether implementations are functionally correct.
- Hands-on proficiency with Linux/Ubuntu, terminal workflows, Git, package managers, test runners, and other developer tooling used to inspect and validate coding tasks.
- Familiarity with coding-agent workflows and tools such as Claude Code, Cursor, OpenCode, or similar, with the ability to review model trajectories and consistently evaluate AI-generated coding solutions.
Perks of Freelancing With Turing
- Work on cutting-edge AI projects with leading foundation model companies
- Collaborate on high-impact work at the frontier of LLM evaluation and reasoning
- Remote, flexible opportunities with global teams
Offer Details
- Commitments Required: 8 hours per day with a 4-hour overlap with PST.
- Employment Type: Contractor position (Note: this role does not include medical/paid leave).
- Duration of Contract: 2 months [expected start date is next week].
Evaluation process
- Complete the take-home assessemnt provided in the Job interest form as a Google Form link. It will take approximately two hours to complete the assessment