freehire launches on Product Hunt on 26 August.

Follow →

Senior AI Backend Engineer — Agent Evaluation & Quality

Summary

Own the evaluation stack for multi-agent systems: design LLM-as-judge components, calibrate against human labels, and quantify agent quality per failure mode.

Salla is seeking a senior engineer to own the evaluation stack for its production multi-agent systems. You will design LLM-as-judge components, calibrate against human labels, and quantify agent quality per failure mode.

You will also build simulators, integrate regression detection into CI, and contribute to agent development by turning failures into improvements. This role emphasizes strong software engineering and production readiness.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available