Software Engineer Pretraining
Summary
At Cursor (Anysphere), this engineer builds the systems that turn raw web data into training-ready datasets for frontier model pretraining — high-throughput data pipelines, data-quality classification models, platform tooling, and web crawling/parsing infrastructure.
You will build systems that transform raw data into training-ready datasets for frontier model pretraining. Depending on the focus area, you will develop data-quality pipelines and models, data-platform tooling, or web crawling and parsing infrastructure. You will make pipelines reliable, observable, reproducible, and scalable.
Responsibilities
- Build high-throughput data pipelines with end-to-end traceability
- Train and ship models that classify rank filter clean and identify data
- Design and run data-mixture repeatability and quality experiments
- Build platforms that transform raw data into training-ready datasets
- Create signals for data quality lineage freshness and pipeline health
- Build and scale web crawling systems
- Improve URL seeding scoring host scheduling crawl success and parsing quality
- Debug and harden crawl infrastructure and automate dataset delivery
Requirements
- Infrastructure or data platform background
- Ability to architect and ship end-to-end with high ownership
- Ability to debug complex systems independently
- Understanding of large-scale distributed systems
- Interest in pretraining data and model quality
- Crawling or search infrastructure experience is a plus
