freehire launches on Product Hunt on 26 August.

Follow →

AI Evaluation Engineer - Proofline

Summary

Design and maintain evaluation harnesses and test infrastructure for AI-facing features, ensuring reliability across different AI assistants and continuous release cycles.

About Us

Meraki Labs (founded by Mukesh Bansal & Peeyush Ranjan) builds and rapidly scales AI-first, "moonshot" startups. We're looking for a high-velocity, production-grade engineer to build 0-to-1 products alongside founders.

About the Role

Proofline is live in production with institutional pilots underway. Building with AI is easy to prototype, but proving reliability in production is a major challenge. In AI-native codebases, verification is the key bottleneck for scaling capabilities. As an AI Evaluation Engineer, you will take ownership of the evaluation scaffolding and quality layer, including eval harnesses for AI-facing features, test infrastructure, and release gates. You’ll play a critical role in monitoring how our product behaves across AI assistants (e.g., Claude, ChatGPT), accounting for differences by host and continual changes.


AI Evaluation Engineer — Evals & Systems Verification

Location: HSR, Bengaluru (On-site)

Experience: 5+ years

Responsibilities

  • Build and maintain evaluation harnesses for AI-facing features to measure and tune system quality (e.g., capture quality, retrieval quality, guidance quality)

  • Own end-to-end and API test infrastructure (Playwright-class), supporting a continuous, daily-release cycle

  • Design and execute host-behavior probes using scripted user sessions across diverse AI assistants, ensuring product behavior aligns with expectations

  • Gate production releases through thorough user acceptance testing (UAT), regression analysis, and quality reporting

Requirements

  • Background in SDET/QA automation or ML evaluation, with proven ownership of test or evaluation infrastructure and not just executing tests

  • Strong programming skills in Python or TypeScript, with experience in API-level testing and tools like Playwright or Cypress

  • Familiarity with LLM applications or a demonstrated interest in evaluating non-deterministic systems

  • Highly autonomous and able to define your own workflows and processes for evaluation and verification


How We Work

Small team, high trust, written decisions. Designs get adversarial review before code; PRs get automated review driven to zero open findings; features aren't done until verified on a live system. AI agents do a large share of the implementation—your leverage is judgment: framing the problem, freezing the right design, and knowing when the machine is wrong.

You Should Apply If

• You enjoy turning messy problems into concrete solutions.

• You thrive in fast-moving, low-process environments.

• You're excited about production engineering (monitoring, reliability, cost, latency).

You Should Not Apply If

• You want remote/hybrid (this is onsite Bangalore only).

• You prefer narrow tickets and minimal ambiguity.

• You prefer highly structured, slow-moving product organizations.

The Recruiter for this role is Ritesh Kalvellu. Apply to the role directly, and he will get back to you if your profile is relevant to the hiring requirements.

What this application asks

ashby

Name, Email, Resume

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available