Software Engineer - Python/Typescript
Summary
Contractor role at Turing evaluating AI coding models: you review AI-generated code and agent behavior in real-world repositories, identify failure modes, and create rubrics and evaluation data to improve model performance. Core stack is Python, TypeScript/JavaScript, or Go, plus LLM-based tooling.
About Turing
About the Role
What You’ll Do
- Evaluate AI-generated code and solutions across real-world software repositories
- Review agent behavior, tool usage, and code changes for correctness and quality
- Identify technical errors, weak approaches, and recurring model failure modes
- Compare model outputs and explain why one solution is better than another
- Create and refine rubrics and evaluation criteria for coding tasks
- Produce high-quality evaluation and preference data used to improve coding models
- Build and maintain pipelines and infrastructure supporting data generation, collection, and evaluation workflows
- Synthesize findings from data work into clear write-ups, updates, and recommendations for the team
- Collaborate closely with researchers and engineers to translate qualitative judgment into scalable processes
- Share clear, actionable findings with AI researchers and engineers
What We’re Looking For
- 5+ years of hands-on software engineering experience
- Strong proficiency in Python, TypeScript/JavaScript, Go, or another major production language
- Experience working in substantial real-world codebases
- Strong code-review skills and technical judgment
- Ability to clearly explain why an implementation is correct, incorrect, or could be improved
- Strong written communication
- Experience using modern LLMs or AI coding tools
Engagement Details
- Compensation: Market rate; please provide a specific hourly rate expectation
- Availability: 40 hours/week preferred, with at least 6 hours of Pacific Time overlap
- Type: Independent contractor
- Duration: Approximately 3 months
- Start: As soon as possible
Evaluation Process
- AI interview (~25 minutes)
- Practical code/AI evaluation exercise (~30 minutes)
- Hiring manager interview (~20 minutes)