Senior Data Engineer
Summary
Emagine is hiring a Senior Data Engineer for a Labeling Data Platform team that handles regulated pharma eLabel content. The role designs, builds and operates scalable pipelines in an Azure Databricks Lakehouse using Python/PySpark, Spark SQL, Delta Lake, medallion architecture, data quality, governance and CI/CD.
Working mode: remote
Contract: B2B
We are looking for an experienced Senior Data Engineer to join our Labeling Data Platform team. The platform supports the collection, processing, transformation, governance, and distribution of labeling and eLabel content across multiple systems and regulatory processes.
The role will be responsible for designing, building, and operating scalable data pipelines and data products within an Azure Databricks Lakehouse environment. You will work closely with business stakeholders, labeling specialists, solution architects, developers, and data consumers to ensure high-quality, reliable, auditable, and compliant data flows supporting global labeling operations.
The ideal candidate combines strong hands-on engineering capabilities with a deep understanding of modern Azure data platforms, Databricks Lakehouse architecture, and regulated data environments.
Your key responsibilities include the following
Develop and maintain data pipelines using Azure Databricks, Python, PySpark, Spark SQL, and Delta Lake.
Design and implement solutions following the Bronze → Silver → Gold medallion architecture.
Ingest and process data from APIs, JSON files, Azure Storage, and enterprise systems.
Build and maintain data quality, validation, reconciliation, monitoring, and automated data testing frameworks, including data quality rules, schema validation, and operational controls
Optimize performance, partitioning, and scalability of data workloads.
Implement CI/CD, automated testing, error handling, and reprocessing processes.
Ensure traceability, auditability, reproducibility, and controlled releases for regulated data pipelines.
Required Qualifications
Strong Azure Databricks engineering with Python/PySpark, Spark SQL and Delta Lake.
Deep knowledge of the Bronze → Silver → Gold medallion architecture shown in the diagram.
Ingestion of structured JSON/API/file data. ADLS Gen2 and Azure storage.
Databricks Workflows/Jobs and preferably Delta Live Tables/Lakeflow where appropriate.
Schema enforcement/evolution. Data-quality rules, validation and reconciliation.
Data transformation and standardization. Metadata management and lineage.
Unity Catalog/governance experience.
API-based ingestion/export patterns.
Handling document metadata and references to PDF/eLabel assets.
Performance optimization and partitioning.
CI/CD for notebooks/code and automated data tests.
Monitoring/error handling/reprocessing.
Experience supporting regulated data pipelines with traceability, auditability, reproducibility and controlled releases.
Experience working in regulated environments with strong compliance and quality requirements.