UC
Unknown company
Senior Prompt Engineer - Evaluation, Scoring and Reliability
Summary
Designs and refines LLM evaluation workflows to score outputs reliably, focusing on semantic alignment, rubric decomposition, and bias mitigation.
We need an experienced prompt engineer to design reliable LLM-based evaluation and scoring workflows. You must understand semantic alignment, rubric decomposition, structured outputs, calibration, prompt sensitivity, model bias, context handling, and failure analysis - not simply write persuasive instructions. The work requires improving scoring consistency across paired inputs, distinguishing genuine meaning from superficial language overlap, and testing performance against paraphrases, contra…