AI Engineer – LLM Data
About the Institute of Foundation Models
We are a dedicated research lab for building, understanding, using, and risk-managing foundation models. Our mandate is to advance research, nurture the next generation of AI builders, and drive transformative contributions to a knowledge-driven economy.
As part of our team, you’ll have the opportunity to work on the core of cutting-edge foundation model training, alongside world-class researchers, data scientists, and engineers, tackling the most fundamental and impactful challenges in AI development. You will participate in the development of groundbreaking AI solutions that have the potential to reshape entire industries. Strategic and innovative problem-solving skills will be instrumental in establishing MBZUAI as a global hub for high-performance computing in deep learning, driving impactful discoveries that inspire the next generation of AI pioneers.
The Role
As an AI Engineer specializing in LLM data, you will build and improve high-quality training data for foundation models across pre-training, mid-training, and post-training. Your work will include large-scale data curation and processing, LLM-based data synthesis, data quality evaluation, and experimentation to understand how different data choices affect model performance.
Key Responsibilities
-
Build, curate, and improve large-scale datasets for LLM pre-training, mid-training, and post-training.
-
Rapidly support time-sensitive data and model development tasks in a fast-moving research environment.
-
Develop and improve data processing pipelines including data extraction, cleaning, filtering, deduplication, quality scoring, transformation, and dataset composition.
-
Design and implement LLM-based data synthesis and augmentation pipelines, including prompt-based generation, filtering, refinement, and quality control of synthetic data.
-
Research and apply methods for improving training data quality, diversity, coverage, and efficiency.
-
Design experiments to understand the relationship between training data and model performance, and use model evaluation results to guide data improvements.
-
Develop scalable tools and workflows for processing and analyzing large datasets efficiently.
-
Analyze datasets using both statistical and model-based methods to identify quality issues, biases, duplication, distributional gaps, and opportunities for improvement.
-
Collaborate closely with researchers, model engineers, and other data teams to translate model development needs into effective data solutions.
-
Document datasets, experiments, data processing methodologies, and key findings clearly to support reproducibility and knowledge sharing.
Professional Experience - Required
- Strong programming skills in Python and experience building reliable engineering or research workflows.
- Hands-on experience with machine learning, deep learning, NLP, or large language models.
- Good understanding of modern LLM development, including how training data is used in pre-training and/or post-training.
- Experience working with large-scale datasets, including data processing, analysis, filtering, transformation, and quality control.
- Ability to independently investigate data or model quality issues, design experiments, and make data-driven technical decisions.
- Familiarity with common machine learning and LLM tools and frameworks such as PyTorch, Hugging Face, vLLM, or similar technologies.
- Strong problem-solving skills and the ability to work effectively in a fast-paced AI research and engineering environment.
- Strong communication and collaboration skills, with the ability to work closely with researchers and engineers across different technical areas.
Professional Experience - Preferred
-
Experience preparing data for large-scale LLM pre-training, continued/mid-training, supervised fine-tuning, preference optimization, reinforcement learning, or other post-training workflows.
-
Experience with synthetic data generation using LLMs, including generation, filtering, verification, or quality evaluation.
-
Experience designing or running LLM evaluations, benchmarks, model training, or fine-tuning experiments.
-
Understanding of how data quality, mixture, diversity, and scaling affect foundation model performance.
-
Experience with large-scale or distributed data processing and compute infrastructure.
-
Experience working with research teams on rapidly evolving foundation model or generative AI projects.
-
Contributions to open-source AI/ML projects, relevant publications, or demonstrated hands-on work with modern foundation models are a plus.
Skills
As published by lever
Resume/CV, Full name, Email, Phone, Current location, Current company, LinkedIn URL, GitHub URL, Personal Website URL, Google Scholar URL, Other URL