freehire launches on Product Hunt on 26 August.

Follow →

Unknown company

Discussion

Senior Prompt Engineer - Evaluation, Scoring and Reliability

Summary

Designs and refines LLM evaluation workflows to score outputs reliably, focusing on semantic alignment, rubric decomposition, and bias mitigation.

We need an experienced prompt engineer to design reliable LLM-based evaluation and scoring workflows. You must understand semantic alignment, rubric decomposition, structured outputs, calibration, prompt sensitivity, model bias, context handling, and failure analysis - not simply write persuasive instructions. The work requires improving scoring consistency across paired inputs, distinguishing genuine meaning from superficial language overlap, and testing performance against paraphrases, contra…

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available