Data Engineer: Multimodal ML Data Pipelines
Summary
Build and maintain scalable data pipelines to clean, version, and serve terabytes of training data for autonomous defense ML models.
Harmattan AI is expanding its data engineering team in Paris, Lausanne, or Zurich to own the data layer powering our autonomous defense ML models. You will manage terabytes of raw data, turning it into clean, versioned training-ready datasets so modelers can focus on model design and training.
You will drive ingestion, curation, labeling workflows, dataset construction, and data lineage to ensure reproducibility and efficient model training in a high-responsibility environment.