Databricks Data Quality & Remediation Architect / Lead Developer - Abu Dhabi
NewBe an early applicant
Databricks Data Quality & Remediation Architect / Lead Developer
We are seeking an experienced Databricks Data Quality & Remediation Architect / Lead Developer to design and build an enterprise-scale Data Quality Rules and Remediation Engine within a Databricks Lakehouse environment with existing team of developers ensuring the team is delivering to plan.
The successful candidate will architect and led the development of a metadata-driven platform that:
Ingests data from enterprise source systems including SAP S/4HANA, SAP BW, Oracle Fusion ERP, Salesforce, ServiceNow and other operational platforms used by government departments.
Profiles, validates and monitors CDEs (critical business entities).
Executes business and technical data quality rules (KPIs).
Identifies data quality defects and potential root causes.
Generates remediation recommendations.
Develops automated correction and enrichment scripts.
Supports working with department source systems leaders to remediate in source applications.
Provides full auditability, governance, lineage and operational monitoring.
The role combines solution architecture, data engineering, data governance, source system integration and software development capabilities.
Key Responsibilities
Solution Architecture
Design and implement a scalable enterprise Data Quality and Remediation Engine on Databricks.
Define the architecture across:
Data ingestion
Profiling
Rule execution
Exception management
Root cause analysis
Automated remediation
Monitoring and reporting
Establish a metadata-driven framework allowing business users and data stewards to configure quality rules without code changes.
Define Bronze, Silver and Gold quality processing layers.
Data Quality Framework Development
Design and develop a reusable rule framework supporting (examples)
Completeness
Mandatory field validation, Null value detection, Missing master records
Accuracy
Business rule validation, Reference data validation
Consistency
Master data synchronisation validation
Validity
Format validation, Legal value checks, Pattern matching (optional)
Uniqueness
Duplicate detection, Fuzzy matching, Golden record identification
Timeliness
Latency monitoring (optional)
Build reusable validation services using:
Databricks SQL
PySpark
Delta Live Tables / Lakeflow
Delta Lake
Unity Catalog
Automated Remediation Development
Design and build automated remediation services including:
Data Correction
Standardisation
Data cleansing
Data enrichment
Format corrections
Reference data alignment
Intelligent Remediation
Pattern-based corrections
AI-assisted recommendations
Duplicate resolution
Master record consolidation
Source System Remediation Scripts
Develop and maintain:
SAP correction scripts
Oracle Fusion correction scripts
Bulk update utilities
Data migration routines
API-based correction services
Ensure all remediation activities include:
Approval workflows
Audit logs
Rollback capability
Change tracking
Segregation of duties controls
Exception Management
Design and implement:
Failed-record quarantine tables
Exception workflows
Root cause categorisation
Issue tracking integration
Remediation queues
Capture:
Rule violated
Business impact
Source application
Affected business object
Recommended action
Resolution status
Data Governance Integration with Microsoft Purview
Collaborate with:
Data Governance teams
Data Stewards
Business SMEs
Support:
Critical Data Elements (CDEs)
Data ownership models
Data quality operating model (tbc)
Stewardship workflows
Dashboards
Data Quality Index (DQI)
Rule pass/fail trends
Exception volumes
Data stewardship actions
Technical Skills
Databricks
Databricks Lakehouse, Delta Lake, Unity Catalog, Databricks Workflows, Lakeflow, Databricks SQL, MLflow, Mosaic AI, Databricks Asset Bundles
Engineering
Python, PySpark, etc
Cloud
Azure Databricks, Azure Data Factory, ADLS Gen2, Azure Key Vault, Azure DevOps