Lead AI Research Engineer — LLM Evaluation & Agent Systems
Summary
The Lead AI Research Engineer will design and implement evaluation frameworks for LLM and agent systems to ensure reliability and performance. This role involves conducting empirical research, building evaluation infrastructure, and guiding technical decisions for AI-native research products.
About the Role
FUNDA is building AI-native research systems for professional users. We are looking for a Lead AI Research Engineer to establish the scientific and technical foundations for evaluating our LLM and agent systems.
This role operates at the intersection of applied AI research, experimental science, and production system architecture. The successful candidate will define how FUNDA measures agent intelligence, reliability, and progress, and will use rigorous empirical evidence to guide decisions across models, prompts, tools, data, and agent architecture.
The role requires more than implementing evaluation pipelines. It requires identifying the right research questions, designing defensible experiments, developing new evaluation methods where established benchmarks are insufficient, and translating findings into company-level technical decisions.
Key Responsibilities
- Define FUNDA’s evaluation strategy, technical architecture, and research methodology for production LLM and agent systems.
- Design evaluation frameworks that faithfully represent real agent behavior, including multi-step reasoning, tool use, information retrieval, long-running execution, and generated artifacts.
- Develop benchmarks that measure not only final-answer quality, but also reasoning trajectories, evidence use, task completion, robustness, and operational reliability.
- Create advanced evaluation methods combining deterministic verification, task-specific grading, process supervision, model-based evaluation, statistical analysis, and targeted expert review.
- Lead controlled experiments and comparative studies across models, prompts, tools, data strategies, and agent architectures.
- Establish rigorous standards for dataset construction, baseline selection, error taxonomy, experiment reproducibility, statistical validity, and regression detection.
- Diagnose complex system failures and distinguish whether they originate from model capability, context construction, tool behavior, data quality, orchestration logic, or infrastructure.
- Develop methods for identifying benchmark contamination, evaluator bias, unstable metrics, and differences between offline evaluation and production performance.
- Translate research findings into architectural recommendations, release criteria, and prioritised engineering investments.
- Build reusable evaluation infrastructure capable of supporting reproducible experiments across different execution environments and agent configurations.
- Set technical standards and mentor engineers in evaluation-driven AI development, experimental reasoning, and evidence-based decision-making.
- Track relevant advances in LLM evaluation, agent benchmarking, AI reliability, and model supervision, and determine which methods are suitable for production adoption.
Expected Impact
The successful candidate will:
- Establish a company-wide evaluation foundation for FUNDA’s core AI systems.
- Create trusted benchmarks that guide model selection, architecture changes, and product releases.
- Enable leadership and engineering teams to distinguish genuine capability improvements from benchmark noise or isolated examples.
- Introduce measurable quality and reliability gates for production agent changes.
- Identify systemic limitations in existing models and agent architectures and propose research-backed solutions.
- Shorten the cycle from discovering a production failure to understanding its cause, validating a solution, and preventing recurrence.
Required Qualifications
- Min Master's Degree in Computer Science, Artificial Intelligence, Machine Learning, Statistics, or a related quantitative field, or equivalent practical expertise.
- A strong track record of conducting applied research and building sophisticated LLM, agent, or machine-learning evaluation systems.
- Deep understanding of experimental design, benchmark construction, statistical reasoning, error analysis, measurement validity, and reproducibility.
- Strong software and systems engineering ability, with experience translating research ideas into reliable production infrastructure.
- Demonstrated ability to evaluate probabilistic systems where correctness cannot be captured by a single metric or static test set.
- Experience analyzing multi-component AI systems and isolating failures across models, data, tools, orchestration, and infrastructure.
- Ability to formulate ambiguous product or research questions as testable hypotheses and measurable evaluation tasks.
- Ability to influence senior technical decisions through clear reasoning, empirical evidence, and well-designed experiments.
Preferred Qualifications
- Experience researching or building tool-using agents, retrieval systems, long-context models, agent memory, or multi-step reasoning systems.
- Experience with process supervision, custom evaluators, model-based grading, human evaluation, adversarial testing, or evaluator calibration.
- Familiarity with AI observability, distributed execution, production experimentation, and large-scale evaluation-data pipelines.
- Experience evaluating AI systems in high-accuracy professional domains such as financial research, scientific research, legal analysis, or enterprise decision support.
- Evidence of technical leadership through research publications, benchmark development, open-source work, patents, technical standards, or significant internal research programs.