Senior Data Engineer (Databricks & Python)
Summary
ITConnect is hiring a Senior Data Engineer in Warsaw to design, build, and operate scalable data pipelines and data products in an Azure Databricks Lakehouse for a labeling/eLabel data platform. Day-to-day work centers on Python, PySpark, Spark SQL, Delta Lake, medallion architecture, data quality, governance, and CI/CD in a regulated environment.
We are looking for an experienced Senior Data Engineer to join our Labeling Data Platform team. The platform supports the collection, processing, transformation, governance, and distribution of labeling and eLabel content across multiple systems and regulatory processes.
The role will be responsible for designing, building, and operating scalable data pipelines and data products within an Azure Databricks Lakehouse environment. You will work closely with business stakeholders, labeling specialists, solution architects, developers, and data consumers to ensure high-quality, reliable, auditable, and compliant data flows supporting global labeling operations.
The ideal candidate combines strong hands‑on engineering capabilities with a deep understanding of modern Azure data platforms, Databricks Lakehouse architecture, and regulated data environments.
- Develop and maintain data pipelines using Azure Databricks, Python, PySpark, Spark SQL, and Delta Lake.
- Design and implement solutions following the Bronze Silver Gold medallion architecture.
- Ingest and process data from APIs, JSON files, Azure Storage, and enterprise systems.
- Build and maintain data quality, validation, reconciliation, monitoring, and automated data testing frameworks, including data quality – rules, schema validation, and operational controls
- Optimize performance, partitioning, and scalability of data workloads.
- Implement CI/CD, automated testing, error handling, and reprocessing processes.
- Ensure traceability, auditability, reproducibility, and controlled releases for regulated data pipelines.
- Strong Azure Databricks engineering with Python/PySpark, Spark SQL and Delta Lake.
- Deep knowledge of the Bronze Silver Gold medallion architecture shown in the diagram.
- Ingestion of structured JSON/API/file data. ADLS Gen2 and Azure storage.
- Databricks Workflows/Jobs and preferably Delta Live Tables/Lakeflow where appropriate.
- Schema enforcement/evolution. Data-quality rules, validation and reconciliation.
- Data transformation and standardization. Metadata management and lineage.
- Unity Catalog/governance experience.
- API-based ingestion/export patterns.
- Handling document metadata and references to PDF/eLabel assets.
- Performance optimization and partitioning.
- CI/CD for notebooks/code and automated data tests.
- Experience supporting regulated data pipelines with traceability, auditability, reproducibility and controlled releases.
- Experience working in regulated environments with strong compliance and quality requirements.