AI Evaluation Engineer - QA, Metrics & Benchmarking
Summary
Own evaluation coverage for Agentic AI systems, focusing on LLM-judge metrics, benchmark dataset curation, and error analysis to support release-readiness for voice and chat solutions.
United States Digital Space LLC seeks an AI Evaluation Engineer to join our AI Evaluation team in Canada. You will own evaluation coverage for the company’s Agentic AI systems alongside the evaluation lead, focusing on LLM-judge metrics, scenario and benchmark dataset curation, and error analysis to support release-readiness for voice and chat solutions.
This role reports to the AI Evaluation manager and may be based in our Vancouver office.