Data Engineer — Data Lakehouse (Hybrid, Kraków, Poland)
Summary
Senior data engineer owning a large-scale data lakehouse in Kraków (hybrid, 2 days/week in office): building bronze/silver/gold layers with Iceberg/Delta tables, CDC streaming ingestion via Kafka/Debezium, query optimization on Trino/Snowflake, and governance plus reconciliation of financial wallet/ledger data. Stack includes Python, Airflow/Dagster, and S3/GCS.
What you'll be doing
Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.
Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.
Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.
Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.
Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.
Monitor ingestion health, data anomalies, and cloud storage/compute spend.
Must-have
Senior: 5+ years in data engineering, with real ownership of a large-scale data lake or lakehouse.
Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.
Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.
CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.
Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from. Critical for financial ledgers.
Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.
Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.
Cloud object storage — S3 or GCS, plus storage tiering and archival.
Python and an orchestrator — Airflow or Dagster, as tools.
Location & work model
Kraków, Poland. Hybrid — 2 days per week from the office.