Software Engineer Agent Evaluation and Quality
Summary
Software engineer at Cursor (Anysphere) building the measurement and evaluation infrastructure for AI agents: creating datasets, replay and scoring systems, dashboards, reliability alerts, and analysis tooling that turn real user signals and agent behavior into product improvements.
You will build measurement, evaluation, and feedback-loop infrastructure for AI agents. You will create datasets, replay and scoring systems, dashboards, reliability alerts, and analysis tooling to turn user signals and agent behavior into product improvements.
Responsibilities
- Design and build AI evaluation systems
- Build feedback loops from real user usage
- Develop tooling to analyze and debug agent behavior
- Define quality metrics, alerts, and triage primitives
- Collect, clean, and interpret user signals
Requirements
- Experience building evaluation or measurement systems
- Data analysis skills
- Ability to collaborate with data scientists and researchers
- Knowledge of model and agent behavior
- Strong software engineering fundamentals

