Agent Evaluation & Evolution Machine Learning Engineer Intern (AML-Ark-US) - 2027 Summer
Summary
Develops and refines automated evaluation systems for LLM agents by analyzing execution traces, user feedback, and designing benchmarks to improve AI agent performance in production environments.
- Analyze agent execution traces and user feedback
- Bring research methods into production
- Build benchmarks and automated judging pipelines
- Design evaluation systems for LLM agents
- Improve systems using feedback loops
Perks/Benefits:
- Health insurance
- Housing allowance
- Life insurance
- Paid Holidays
- Paid sick time
- Wellbeing benefits