freehire launches on Product Hunt on 26 August.

Follow →

Member of Technical Staff (Evals)

Summary

As a Member of Technical Staff on Evals, you'll own measuring AI Employee performance via evaluation systems, feedback loops, and autonomous failure analysis.

As a Member of Technical Staff on Evals, you'll own how we measure AI Employee performance on the job.

Areas you may work in:

  • Evaluation systems: datasets, offline replay, scorers, and regression alerts
  • Feedback loops for models and agents
  • Autonomous failure analysis
  • Defining what good, degraded, and failed runs look like, and alerting on them

You may be a good fit if you:

  • Have built evaluation or measurement systems, such as AI evals, experimentation, ranking, or search quality
  • Can turn ambiguous quality questions into concrete metrics and decisions
  • Have strong software engineering fundamentals and ship production systems
  • Are comfortable where there are few established patterns
  • Want to work in person in San Francisco

Even better:

  • You've shipped LLM systems to production and learned where they break
  • You have strong data instincts and work well with researchers
  • You contribute to open source projects

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available