freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineering

Summary

Build and own a Databricks Lakehouse data warehouse and real-time pipelines for a financial trading business, modeling customer, order, trade, and market data with Kafka, Flink, and Spark.

Responsibilities



  • Own the data-warehouse architecture and ELT pipelines for financial trading business; build a layered model (ODS / DWD / DWS / ADS) on Databricks Lakehouse.

  • Lead subject-area modeling for customer, account, order, trade, position, clearing & settlement and market data; productize reusable metrics and customer profiles.

  • Build the real-time computing stack on Kafka + Flink / Spark Structured Streaming - CDC ingestion and low-latency pipelines powering real-time trade monitoring and live metrics.

  • Drive the migration from classic batch warehousing to a lakehouse (Delta Lake / Iceberg) - unified streaming + batch writes, consistency, rollback and schema evolution.

  • Define and operate the bar for data quality, metadata, lineage, SLA and change-management; guard warehouse and pipeline (incl. streaming) reliability.

  • Partner deeply with trading, product, finance and compliance to land a single source of truth and self-service consumption.


Requirements



  • Bachelor's or above in CS / Software / Math / Statistics; 8+ years in DW / big-data engineering.

  • Solid financial-industry background: 3+ years building warehouses in securities, exchanges, brokerage, payments or banking; exchange-business data experience (orders, trades, positions, clearing & settlement, market data) preferred.

  • Expert in Kimball dimensional modeling, SCD and metric-layer design.

  • Hands-on with Databricks (Delta Lake, Unity Catalog, Spark SQL / PySpark, DLT).

  • Production experience in real-time computing: expert in Flink or Spark Structured Streaming, with Kafka, CDC (e.g. Debezium), exactly-once semantics and state management.

  • Familiar with mainstream open-source big-data frameworks in production - Spark / Flink / Kafka / Iceberg / Hudi / Airflow / DolphinScheduler - able to make technology choices and debug at the framework level.

  • Deep SQL & Spark tuning - partitioning, Z-Order, shuffle, broadcast, AQE.

  • Familiar with governance tooling (DataHub / Atlas / Data Map).

  • Strong cross-functional, upward-communication and written-spec skills.

  • Able to use chinese and english as daily working language to work wit chinese speaking stakeholders

See also