Data Engineer for Generative AI Data Lakes (Hybrid/Remote)
Summary
Build and maintain scalable data pipelines and lakes for a generative AI video model, ensuring high-quality annotated datasets for model training and feature development.
Synthesia is seeking a Data role at the intersection of applied research, data engineering, and ML infrastructure. You will manage the data lifecycle for researchers, working with over a million hours of video and audio data to power model improvements and new features.
The role emphasizes data quality, annotation, and scalable pipelines in a hybrid Europe setting. You will collaborate closely with model training teams to build a world-class data lake and drive longer-term data strategy.