Point your AI agent at freehire and let it find you a job.

Get the CLI →

Union Rebellion

NewBe an early applicant

Software Engineering, Data Science & Systems Design Experts (AI Model Evaluation)

Posted Updated 1 view
Discussion
Role Overview

As a Technical AI Evaluator, you will assess and improve AI-generated responses to coding and software engineering tasks. Your work will directly impact how accurately and effectively AI systems solve problems, write code, and communicate technical concepts.

This is a hands-on, intellectually rigorous role suited for professionals with strong coding, analytical, and problem-solving skills.

Key Responsibilities
  • Evaluate AI-generated responses for accuracy, reasoning, clarity, and completeness

  • Execute and test code to validate correctness and performance

  • Identify bugs, inefficiencies, and logical flaws in model outputs

  • Annotate responses with detailed feedback and improvement suggestions

  • Assess code quality, readability, and adherence to best practices

  • Fact-check technical claims using reliable and authoritative sources

  • Apply structured evaluation frameworks and guidelines consistently

What We’re Looking For
  • Bachelor’s, Master’s, or PhD in Computer Science or a related field

  • 3+ years of professional experience in software engineering, data science, or related roles

  • Strong proficiency in at least two programming languages (e.g., Python, JavaScript, C++, Java, SQL, Go, Rust, etc.)

  • Ability to independently solve medium to hard algorithmic problems

  • Experience working with or alongside LLMs in coding workflows

  • Strong attention to detail and ability to evaluate complex technical reasoning

Nice to Have
  • Experience with AI model evaluation, RLHF, or data annotation

  • Background in competitive programming

  • Contributions to open-source projects (merged pull requests preferred)

  • Experience reviewing production-level code

  • Ability to explain complex technical concepts clearly to non-experts

What Success Looks Like
  • You consistently identify subtle bugs, flawed logic, and edge cases

  • Your feedback improves the accuracy and clarity of AI-generated code

  • You produce structured, reproducible evaluation outputs

  • AI systems become more reliable for real-world engineering tasks

Compensation & Terms
  • Independent contractor engagement

  • Fully remote with flexible working hours (open to US and non-US candidates)

  • Full-time or part-time contract options available

  • Weekly payments based on services rendered

Equal Opportunity

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

See also

Data Science jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available