AI Test Engineer
Summary
An AI Test Engineer at 01.AI in Beijing will own testing and evaluation for enterprise AI Agent projects — defining test strategies, building evaluation datasets, benchmarks and harnesses, and testing knowledge retrieval, tool calls, multi-turn conversations, and workflows. Core work combines Python-based tooling with human, automated, and model-assisted evaluation, plus stability, performance, se
You will own testing and evaluation delivery for enterprise Agent projects. You will define test strategies, evaluation plans, and acceptance criteria; identify quality risks; and provide launch, delivery, and customer-acceptance recommendations. You will evaluate AI applications, knowledge retrieval, Agents, tool calls, multi-turn conversations, and workflows. You will build enterprise evaluation systems, datasets, benchmarks, regression mechanisms, Harness capabilities, and evaluation-platform components. You will also lead quality initiatives for stability, performance, security, and adversarial testing.
Responsibilities
- Define test strategies, evaluation plans, and acceptance criteria for enterprise Agent projects
- Identify quality risks from customer business and usage scenarios
- Drive product, algorithm, engineering, and delivery teams to resolve issues
- Deliver evaluation conclusions and improvement recommendations for launch, delivery, and customer acceptance
- Test and evaluate AI applications, knowledge retrieval, Agents, tool calls, multi-turn conversations, and workflows
- Design combined human, automated, and model-assisted evaluation approaches
- Analyze logs and call chains, drive fixes, and verify regressions
- Build quality metrics, evaluation datasets, standards, baselines, release thresholds, and feedback mechanisms
- Build Harness and evaluation-platform capabilities for execution, scoring, comparison, analysis, and reporting
- Lead stability, performance, security, and adversarial testing initiatives
- Develop testing and evaluation scripts, tools, or platform components
Requirements
- 5 or more years of testing, test development, quality, or evaluation experience
- 2 or more years of AI or Agent evaluation experience
- Ability to independently deliver complex projects
- Harness knowledge and practical testing or evaluation Harness experience
- AI and Agent evaluation methods
- Evaluation dataset development, metric design, and combined human and automated evaluation
- Knowledge retrieval, Agent, tool calling, multi-turn conversation, and workflow knowledge
- Python
- Enterprise project experience
- Data isolation, permission management, auditing, release rollback, and private deployment quality requirements
- Cross-team delivery ability