freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer (AI Native, Pretrain Algorithm, data pipeline, production)

Summary

Build and scale data pipelines that process billions of documents for training large language models, optimizing data quality and performance for AI research.

About the Role

Our client is a well-funded AI research company at the frontier of large language model development, building foundational models that power next-generation AI applications. As the team scales its pre-training infrastructure, they are looking for a skilled Data Engineer to own the end-to-end data pipeline — from raw web-scale ingestion through to high-quality tokenized training datasets. This is a high-impact, technically deep role sitting at the intersection of distributed systems engineering and LLM research, and is ideal for engineers who care deeply about data quality and scale.

Key Responsibilities

  • Build and scale data pipelines processing billions of documents for large-scale model training
  • Process and structure web and document data, including text, tables, formulas, and code
  • Develop data filtering, quality evaluation, and deduplication systems using rules, ML models, and LLMs
  • Optimize large-scale data processing for performance, scalability, and cost efficiency
  • Run experiments to evaluate the impact of data strategies on model performance
  • Collaborate with LLM researchers and ML engineers to continuously improve pre-training data quality

Requirements

  • Strong software engineering fundamentals with expert-level Python proficiency and a track record of production-quality code
  • Hands-on experience with at least one distributed processing framework such as Spark, Ray, Flink, or Beam
  • Proven experience designing and data cleaning pipelines, with a solid understanding of cost and throughput trade-offs
  • Deep familiarity with NLP/LLM data concepts including tokenisation, deduplication, text extraction, and quality filtering
  • Experience building pre-training or mid-training datasets from scratch (is a bonus)

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available