Senior AI Backend Engineer — Agent Evaluation & Quality
Summary
Own the evaluation stack for multi-agent systems: design LLM-as-judge components, calibrate against human labels, and quantify agent quality per failure mode.
Salla is seeking a senior engineer to own the evaluation stack for its production multi-agent systems. You will design LLM-as-judge components, calibrate against human labels, and quantify agent quality per failure mode.
You will also build simulators, integrate regression detection into CI, and contribute to agent development by turning failures into improvements. This role emphasizes strong software engineering and production readiness.
#J-18808-Ljbffr