Agent Evaluation & Evolution Machine Learning Engineer Intern (AML-Ark-US) - 2027 Summer
Summary
An internship evaluating and improving AI agent systems by analyzing execution traces, user feedback, and failure patterns to design benchmarks and production-ready evaluation pipelines for LLM agents.
- Analyze agent execution traces
- Analyze user feedback
- Build benchmarks and automated judging pipelines
- Collaborate with research platform and product teams to deploy methods to production
- Design LLM agent evaluation systems
- Improve systems based on failure patterns
- Support experience to capability closed loop
Perks/Benefits:
- Health insurance
- Housing allowance
- Life insurance
- Paid Holidays
- Paid sick time
- Wellbeing benefits