Founding Engineer
Ressl AI Founding Engineer
We are building benchmarks with vertical AI companies. This role would mostly involve building evals. You will spend most of your time deciding whether an AI system actually did the job, then building the datasets, graders, and checks that proves it. We are a small YC team. You work with the founders. If you want a 9-to-5 or a remote-first job, skip this. If you want to own how we measure things, keep reading.
What you will do
- Turn real failures into tasks we can score
- Build datasets, gold labels, graders, and human review loops
- Know when a metric is lying, and say so
- Ship changes to the eval setup
- Find LLM incapabilities in production systems
You will like this if
- You have shipped agents.
- You have built evals, graders, or datasets for LLMs or agents
- You have a story about a metric that lied
- You can sit with a lot of traces and find the few that matter