Point your AI agent at freehire and let it find you a job.

Get the CLI →

CommIT

NewBe an early applicant

Caliente Interactive: Data Engineer — Data Lakehouse Senior

Posted 1 view
Discussion

Summary

Senior Data Engineer managing a large-scale data lakehouse for financial events in the iGaming sector. The role involves designing lakehouse architecture (Iceberg/Delta), implementing CDC streaming with Kafka/Debezium, and ensuring data governance and reconciliation across S3, Snowflake, and Databricks.

We are looking for Data Engineer in Kraków, Poland who will own the company data lake — the system of record for millions of financial events a day (bets, wallet movements, live odds) across 12M+ active users. You decide how that data lands, is stored, retained, and governed on S3 + Snowflake/Databricks, so analytics, finance, and regulators all see accurate, reconciled data with zero drift from source.

Domain: Regulated iGaming / wallet & ledger data. Audit-heavy: regulators, finance and analytics all consume the same tables. Millions of financial events per day, terabyte-plus scale.

What you'll be doing:

  • Own the lakehouse architecture: bronze/silver/gold layers, Iceberg/Delta tables, schema evolution.
  • Land operational data via CDC streaming (Kafka, Debezium), handling late and duplicate events.
  • Design data layout for speed and cost: partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
  • Own retention and archival: storage tiering, regulatory retention, immutability, GDPR deletion.
  • Guarantee correctness: freshness SLAs, drift detection, reconciliation against the source wallet and ledger systems.
  • Own governance: catalog and lineage, row/column access control, PII masking, encryption, audit trails.
  • Monitor ingestion health, data anomalies, and cloud storage/compute spend.


Requirements

Must-have:

  • 5+ years in data engineering, with real ownership of a large-scale data lake or lakehouse.
  • Lakehouse architecture — bronze/silver/gold layering, an open table format (Iceberg, Delta, or Hudi), schema evolution.
  • Data layout & query optimization at TB+ scale — partitioning, compaction, file sizing, query performance on Trino/Athena/Snowflake.
  • Cloud lakehouse/DWH in production — Snowflake, Databricks, or BigQuery.
  • CDC & streaming ingestion — Kafka + Debezium or equivalent; late, duplicate and out-of-order events.
  • Strong SQL and data modeling — enough relational grounding to reason about the OLTP systems you capture from. Critical for financial ledgers.
  • Correctness — freshness SLAs, drift detection, reconciliation against source wallet/ledger systems.
  • Governance — catalogs, lineage, row/column access control, PII masking, retention, GDPR deletion.
  • Cloud object storage — S3 or GCS, plus storage tiering and archival.
  • Python and an orchestrator — Airflow or Dagster, as tools.

Location & work model:

Kraków, Poland. Hybrid — 2 days per week from the office.

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available