Large Language Model Data Engineer
Summary
Builds and optimizes distributed web-crawler and big-data pipelines (scheduling, collection, parsing, storage) that feed Baichuan's LLM and RAG/search systems. Day-to-day work spans crawling tools like Scrapy and Fiddler, anti-bot techniques, and the Hadoop/Spark/Hive/HBase stack using Python, Java, or Go.
You will build and optimize distributed crawler workflows covering scheduling, collection, parsing, and storage. You will develop targeted collection processes and domain knowledge bases, contribute to retrieval-augmented generation and search algorithms, and improve the stability and timeliness of data throughout the pipeline.
Responsibilities
- Build and optimize distributed crawler systems
- Optimize data scheduling, collection, parsing, and storage workflows
- Develop targeted data collection processes and domain knowledge bases
- Contribute to content understanding, retrieval, and ranking algorithms for RAG systems
- Clean, process, and analyze data to improve end-to-end data stability and timeliness
Requirements
- Bachelor's degree or above
- At least 2 years of web crawling or big-data processing experience
- HTTP and TCP knowledge
- Web crawling technology and tools including Fiddler and Scrapy
- Hadoop, Spark, Hive, and HBase
- Python, Java, or Go
- Web crawler anti-bot technology
- APK unpacking and reverse engineering is preferred
- APK or mini-program data collection experience is preferred
- PyTorch or TensorFlow is preferred
- Deep learning, machine learning, and natural language processing knowledge is preferred
- Big-data framework internals and data-processing bottleneck optimization are preferred
- Big-data component operations and maintenance ability is preferred