Senior Data Engineer / ML Platform
Summary
Build and scale a cloud-native data platform on AWS to power real-time recommendations and analytics for Europe’s leading brands.
Sovendus connects Europe’s leading brands with customers who buy.
We’ve been doing it for 17 years, with one clear focus: turning general checkout traffic into a source of new customers and incremental revenue.
More than 3,000 brands across Europe are already growing with Sovendus through long‑term partnerships and real European transaction data. That’s how growth becomes predictable. And scalable.
Now, we’re expanding our international data team — and we’re looking for a Data Engineer based in Spain who wants to take ownership and build at scale.
As a Data Engineer, you will:
- Work in a modern, cloud-only AWS stack — EMR, EKS, S3, Athena/Glue, SageMaker, DynamoDB — orchestrated with Airflow and managed as code with Terraform
- Shape the Kafka backbone that carries user-journey events to the data lake, ML training pipelines, and analytics consumers
- Move data at scale using Spark SQL on Delta Lake and Apache Iceberg, enabling reliable, ACID-compliant data processing on S3
- Build the data layer behind our recommendation engine — from Athena schemas for model training to real-time data assets in Redis and DynamoDB
- Push AI-assisted development and operations — using AI tooling across coding, review, documentation, and platform workflows
- Own the long‑term health of the platform — driving upgrades, reducing tech debt, and keeping the stack modern and scalable
Your profile:
- You have 5+ years of experience in data engineering or ML platform work in production environments
- You write strong Python code, focused on production-grade applications (not just notebooks or scripts)
- You have advanced SQL skills (Athena / Spark SQL) and are comfortable working in SQL-first data environments
- You have hands‑on experience with Kafka (producers, consumers, stream processing)
- You bring solid cloud experience (AWS preferred, or GCP/Azure)
- You are fluent in English and comfortable working in an international team
- You are based in Spain and have valid work authorization
It’s a plus if you have:
- Experience with Spark and modern data lake ecosystems (PySpark, Delta Lake, Iceberg), including job structuring and performance tuning
- Exposure to orchestration and MLOps tools such as Airflow, SageMaker, or similar
- Familiarity with ML data workflows, including feature pipelines, training datasets, and online serving
You’ll be part of an international, fast-growing company that values ownership, trust and collaboration. Your work will directly shape how our network grows in Italy.