Research Crawling Engineer
You will design and operate large-scale web data acquisition systems for research and model development. You will build distributed crawlers, handle anti-bot systems and dynamic websites, develop data processing pipelines, construct research datasets, monitor crawl quality, collaborate with research teams, and optimize infrastructure for cost, latency, and reliability.
Responsibilities
- Build and maintain large-scale web crawlers
- Design high-throughput fault-tolerant data collection systems
- Handle anti-bot systems rate limits and dynamic websites
- Develop pipelines for cleaning deduplication filtering and normalisation
- Construct and maintain datasets for research and model training
- Monitor crawl performance coverage and data quality
- Collaborate with research teams on data collection needs
- Optimize infrastructure for cost latency and reliability
Requirements
- Strong programming experience in Go Rust Python Java or C++
- Experience building web crawlers or large-scale data pipelines
- Understanding of HTTP networking and browser behavior
- Familiarity with distributed systems and parallel processing
- Experience with large datasets at TB to PB scale
- Ability to debug unstable or adversarial environments
Benefits
- Fully remote work
- Benefits package
- Equity package