Point your AI agent at freehire and let it find you a job.

Get the CLI →

Staffy

NewBe an early applicant

QA / AI Evaluation Engineer

Posted
Discussion

Summary

Senior QA engineer specializing in AI evaluation: designs and runs evaluation strategies for AI-powered apps, from human-in-the-loop UAT batches to large automated LLM evaluation suites, building Python-based metrics and testing frameworks to measure factual grounding, accuracy, and regressions, integrated into CI/CD pipelines.

About the company

We are a young and fast-growing recruiting company with five years of experience working across Latin America and the United States. We partner closely with teams and founders to help them build strong, high-impact teams through recruitment, outsourcing, and team-building services.

Our culture is built on effective communication, trust, and transparency. We believe great work happens when people feel heard, supported, and empowered to grow. Today, our team is made up of more than 80 professionals working across different projects throughout the region, collaborating remotely and learning from each other every day.

About the role

We are looking for a Senior QA / AI Evaluation Engineer to join an innovative team focused on measuring and continuously improving the quality, accuracy, and reliability of AI-powered applications.

In this role, you will design and execute AI evaluation strategies at scale, ranging from human-in-the-loop UAT batches to hundreds of thousands or millions of automated evaluations. You will develop the metrics and testing frameworks needed to objectively measure factual grounding, accuracy, quality improvements, and regression across AI-generated responses.

This position combines QA Engineering, test automation, data analysis, and AI evaluation, requiring strong analytical skills and the ability to translate evaluation results into actionable improvements.

Responsibilities

  • Design and execute AI evaluation strategies across large question sets, from human UAT batches to large-scale automated evaluation suites
  • Develop methodologies to statistically measure factual grounding, accuracy, and quality improvements across different versions of AI solutions
  • Build and maintain metrics frameworks to quantify improvements in AI-generated responses and their alignment with source content
  • Design and execute load, performance, quality, and scalability tests as the underlying knowledge corpus and AI workloads grow
  • Define and maintain test plans, test cases, evaluation criteria, and quality gates for AI-powered applications
  • Automate regression and AI evaluation suites and integrate them into CI/CD pipelines
  • Build evaluation harnesses using Python and appropriate data analysis and testing tools
  • Work with multiple LLM providers and evaluate differences in model outputs, accuracy, and grounding
  • Analyze evaluation results and identify patterns, regressions, and opportunities for improvement
  • Collaborate closely with Engineering and AI teams to reproduce issues, validate fixes, and ensure quality improvements
  • Communicate quality metrics, findings, and recommendations clearly to both technical and non-technical stakeholders
  • Continuously improve QA processes, evaluation methodologies, automation coverage, and quality standards

Requirements

  • 4+ years of professional experience in QA, Test Engineering, or a related quality-focused role, with exposure to data, ML, or AI systems
  • Strong Python skills and experience applying data analysis or data science techniques to evaluate system quality
  • Hands-on experience with LLM evaluation, AI testing, or AI quality assessment
  • Experience using statistical analysis to measure accuracy, factual grounding, quality, or performance improvements
  • Ability to build automated, large-scale evaluation and regression harnesses, combined with human-in-the-loop testing approaches
  • Experience with test automation frameworks, scripting, and CI/CD integration
  • Comfortable evaluating and comparing outputs from multiple LLM providers
  • Strong understanding of QA methodologies, test planning, regression testing, quality gates, and defect management
  • Strong analytical and problem-solving skills, with a high level of attention to detail
  • Excellent written and verbal communication skills and the ability to present findings to different audiences
  • Bachelor's degree in a technical field or equivalent practical experience.

Nice to have

  • Experience with LLM evaluation frameworks and AI observability tools
  • Experience designing evaluation systems for RAG, conversational AI, or AI agents
  • Experience with large-scale data processing and statistical analysis
  • Previous experience working in technology consulting or client-facing environments
  • Experience in Life Sciences, Healthcare, or Pharmaceutical industries
  • Familiarity with cloud platforms and modern AI/ML infrastructure
  • Experience designing quality frameworks for systems operating at high scale

Benefits

  • People First culture
  • Referral Program
  • Free access to streaming platforms
  • Free access to Spotify Premium
  • GYM discount
  • Travel discount
  • E-Learning discount
  • Birthday-day gift
  • Points Program

Skills

See also

QA jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available