Senior Data Engineer (GenAI / LLM) AI, GenAI, LLM, SQL, Kafka Warsaw, Gdańsk, Gdynia
Summary
Builds and maintains AI-driven data pipelines and LLM-assisted workflows to automate engineering processes, schema governance, and test automation using Kafka, SQL, and distributed frameworks.
Responsibilities
- Design and implement LLM‑and agent‑assisted workflows to improve developer productivity and reliability of data pipeline delivery
- Build and enhance AI‑driven automation for engineering processes, including: pull request automation, quality gates and assistive development tooling
- Implement AI‑supported schema review and governance gates within the canonical message lifecycle
- Drive AI‑powered test automation, including automated E2E test generation and quality validation
- Develop and maintain data packaging pipelines transforming source data into canonical message formats aligned with Data Packaging Framework and Availability Layer patterns
- Build and manage SQL‑based data transformations and configuration‑driven pipelines
- Support the canonical message lifecycle, including:schema validation, versioning, compatibility checks and publishing workflows
- Contribute to framework releases, operational stability, and documentation
- Maintain and update Architecture Decision Records (ADRs)
Requirements
- Hands‑on experience building LLM‑powered automation for engineering workflows (e.g., review gates, quality checks, assistant tooling for dev/test/docs)
- Experience with AI‑assisted review/approval flows and integrating AI steps into CI/CD‑style pipelines
- Experience with AI‑assisted test generation / test automation concepts
- Strong engineering discipline around quality and safety for AI outputs (evaluation mindset, reproducibility, traceability)
- Strong experience in data engineering and distributed processing; practical delivery mindset
- Hands‑on experience with Kafka and SQL‑based transformations
- Experience with batch/stream processing frameworks (e.g., Flink or Spark)
- Familiarity with schema formats and lifecycle (e.g., JSON/Avro/XSD) and versioning/compatibility concepts
- Experience working with Git workflows and Kubernetes environments
- CI/CD and build pipeline experience (Bazel + Tekton)
Nice to have
- Experience with Data Reference Architecture / data product model
- Knowledge of Airflow, Snowflake, Iceberg, or similar data platforms
- Exposure to data governance, metadata, and lineage concepts
- Experience in banking or other regulated environments
Offer
- Private medical care
- Co‑financing for the sports card
- Constant support of dedicated consultant