Member of Technical Staff - Data Ingestion Engineer
Summary
Builds and operates large-scale data ingestion systems (crawling, extraction, normalization, versioning) that feed Reflection's LLM pre-training, running experiments on data quality and developing specialized crawlers. Core technologies include distributed data frameworks like Ray, Beam, or Spark.
You will build and operate systems that acquire, extract, normalize, version, and deliver large-scale data for pre-training. You will run experiments on crawling and extraction strategies, analyze data quality and coverage, develop specialized crawlers, and improve ingestion infrastructure through code review and production debugging.
Responsibilities
- Build and operate large-scale data ingestion systems for pre-training
- Run experiments on crawling strategies, extraction methods, and ingestion tradeoffs
- Analyze ingested data to identify gaps, redundancy, and improvements
- Build reliable ingestion pipelines for large data campaigns
- Develop specialized crawlers for high-priority data sources
- Review code, debug production issues, and improve ingestion infrastructure
Requirements
- Experience building web crawling, data ingestion, or large-scale data acquisition systems
- Experience with Ray, Beam, Spark, or similar technologies
- Familiarity with LLM training and evaluation
- Experience working with multi-terabyte to petabyte-scale datasets
- Ability to design experiments and use data to improve systems
- Communication skills
Benefits
- Stock options
- Medical, dental, vision, and life insurance
- Annual wellness allowance
- Daily office lunch and dinner
- 22 weeks of paid parental leave
- Unlimited paid time off in the U.S.
- 30 vacation days in the U.K.
- Visa sponsorship support
- Regular off-sites, happy hours, and team celebrations