Data Engineer - Equipo de LLMs Training - IT
Summary
Build and scale data pipelines and tools to prepare and process large-scale text datasets for training large language models.
- Build data ingestion pipelines
- Collaborate with data scientists and ML engineers to build training pipelines
- Develop internal tools for dataset preparation
- Implement data curation quality deduplication scoring diversification
- Process textual data at scale for LLM training
- Scale synthetic data generation with open source and commercial models
Perks/Benefits:
- Hybrid work