freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer (GenAI / LLM) AI, GenAI, LLM, SQL, Kafka Warsaw, Gdańsk, Gdynia

Summary

Builds and maintains AI-driven data pipelines and LLM-assisted workflows to automate engineering processes, schema governance, and test automation using Kafka, SQL, and distributed frameworks.

Responsibilities

  • Design and implement LLM‑and agent‑assisted workflows to improve developer productivity and reliability of data pipeline delivery
  • Build and enhance AI‑driven automation for engineering processes, including: pull request automation, quality gates and assistive development tooling
  • Implement AI‑supported schema review and governance gates within the canonical message lifecycle
  • Drive AI‑powered test automation, including automated E2E test generation and quality validation
  • Develop and maintain data packaging pipelines transforming source data into canonical message formats aligned with Data Packaging Framework and Availability Layer patterns
  • Build and manage SQL‑based data transformations and configuration‑driven pipelines
  • Support the canonical message lifecycle, including:schema validation, versioning, compatibility checks and publishing workflows
  • Contribute to framework releases, operational stability, and documentation
  • Maintain and update Architecture Decision Records (ADRs)

Requirements

  • Hands‑on experience building LLM‑powered automation for engineering workflows (e.g., review gates, quality checks, assistant tooling for dev/test/docs)
  • Experience with AI‑assisted review/approval flows and integrating AI steps into CI/CD‑style pipelines
  • Experience with AI‑assisted test generation / test automation concepts
  • Strong engineering discipline around quality and safety for AI outputs (evaluation mindset, reproducibility, traceability)
  • Strong experience in data engineering and distributed processing; practical delivery mindset
  • Hands‑on experience with Kafka and SQL‑based transformations
  • Experience with batch/stream processing frameworks (e.g., Flink or Spark)
  • Familiarity with schema formats and lifecycle (e.g., JSON/Avro/XSD) and versioning/compatibility concepts
  • Experience working with Git workflows and Kubernetes environments
  • CI/CD and build pipeline experience (Bazel + Tekton)

Nice to have

  • Experience with Data Reference Architecture / data product model
  • Knowledge of Airflow, Snowflake, Iceberg, or similar data platforms
  • Exposure to data governance, metadata, and lineage concepts
  • Experience in banking or other regulated environments

Offer

  • Private medical care
  • Co‑financing for the sports card
  • Constant support of dedicated consultant

See also