Senior Data Engineer
Summary
Senior Data Engineer builds and maintains secure enterprise data pipelines for large-scale digital transformation projects in customs, defense, and healthcare, using Spark, SQL, and modern data architecture.
About the project and the role
We are supporting an important project led by a global company that specializes in digital transformation. The goal is to build a secure enterprise data platform, often compared to a "French Palantir"-style solution.
You will join a small team of 10–15 experienced experts and work directly with project leaders on large Data & AI initiatives across Europe. These projects support areas such as customs, national defense, and healthcare, and are worth many millions of euros.
As a Senior Data Engineer, you will design, build, and maintain data pipelines. Your work will help collect, process, and deliver high-quality data for business intelligence, APIs, data science, and machine learning applications.
Key responsibilities
Design, build, and maintain ETL/ELT data pipelines.
Connect different data sources, including internal systems, APIs, databases, files, IoT devices, and logs.
Develop data processing solutions using Spark, SQL, or similar technologies.
Organize data using Bronze, Silver, and Gold data layers.
Improve pipeline performance, scalability, and infrastructure costs.
Create automated data quality checks for freshness, completeness, and schema validation.
Handle failed data loads, retries, and data reprocessing.
Support data governance by contributing to data contracts and certification rules.
Use Git for version control and automate deployments with CI/CD.
Document datasets and metadata.
Monitor data pipelines, solve incidents, and support business continuity.
Technical requirements
Required
At least 5 years of experience as a Data Engineer.
Strong SQL skills, including query optimization and data modeling.
Experience with distributed processing tools such as Spark, Flink, or similar.
Experience with workflow orchestration tools like Airflow, Dagster, or similar.
Good knowledge of data formats such as Parquet, Avro, and ORC.
Experience with analytical data modeling, including star schemas and wide tables.
Nice to have
Experience with streaming and Change Data Capture (CDC) technologies such as Kafka, Pulsar, or Debezium.
Knowledge of analytics engineering tools like dbt.
Experience with data quality testing and pipeline monitoring.
Good understanding of Git and CI/CD practices.
Additional skills
Ability to balance data volume, processing speed, and infrastructure costs.
Strong understanding of the Data-as-a-Product approach.
Good communication skills in French and English to work with BI teams, Data Scientists, and business stakeholders.
Careful, reliable, and focused on delivering stable production systems.
Key performance indicators (KPIs)
Success in this role will be measured by:
Time needed to deliver usable datasets from raw data.
Pipeline reliability, processing speed, and latency.
Number of data-related incidents.
Processing and infrastructure cost efficiency.
How easily datasets can be reused by different teams.
