Data Platform Engineer: ETL, MDM & AI Pipelines
Summary
Build and maintain governed data pipelines (ETL/ELT, OCR, embeddings) that feed AI systems and KPI dashboards from enterprise sources like ERP and HR.
Build the platform's governed data layer: ETL/ELT pipelines from enterprise systems, data quality and reconciliation, the KPI dictionary and authoritative source register, and the migration and knowledge-ingestion pipelines (OCR, embeddings, indexing) that ground the AI.
Key Responsibilities
- Build ingestion pipelines for the priority systems agreed in discovery, through the approved gateway patterns.
- Run data profiling and quality assessment; implement quality rules, reconciliation and corrections.
- Build and maintain the KPI dictionary, calculations and authoritative data source register.
- Deliver migration and knowledge ingestion per domain within the agreed tiers.
- Evidence data lineage for every critical KPI and agent knowledge source.
Requirements
- 5+ years in data engineering or integration, including ERP/HR/ITSM-class sources.
- Experience with document processing pipelines (OCR, chunking, embeddings) is a strong advantage.
- Government data governance exposure is an advantage.
- ETL/ELT design across APIs, database extracts, file/batch and event feeds.
- Data profiling, quality rules, exception handling and steward correction workflows.
- SQL and pipeline orchestration; master data and entity mapping.
- Data lineage and traceability from source to dashboard and agent answer.
- Data classification, retention and residency compliance practice.