Data Engineer
Summary
The Data Engineer builds and maintains data pipelines and knowledge repositories to prepare structured and unstructured data for Generative AI and RAG-based solutions. The role ensures data accuracy, governance, and quality across enterprise systems to support reliable AI-driven business decisions.
Role Purpose
The Data Engineer prepares the Group s enterprise data and knowledge sources so that AI Generative AI and Agentic AI solutions operate on information that is accurate current governed and fit for purpose Spanning structured data in core business systems and unstructured content across documents policies and correspondence the role builds and maintains the pipelines and knowledge repositories that underpin retrieval-augmented AI It is the control point that ensures agents draw only on trusted approved sources making this role the single greatest determinant of whether the Group s AI outputs can be relied upon in business decisions
Data Preparation for AI Use Cases
Prepare enterprise data and knowledge sources for AI use cases working from prioritised business requirements defined with the AI Agentic AI Lead Support data extraction cleansing classification tagging and indexing across structured semi-structured and unstructured sources Build and maintain ingestion pipelines from ERP CRM HRMS procurement systems the data lake and document repositories Design chunking metadata and enrichment strategies that materially improve retrieval relevance and answer quality Handle multi-format content PDF Office documents scanned material and email including OCR and text extraction where required
Knowledge Repositories RAG Enablement
Build and maintain knowledge repositories for RAG-based AI solutions including embedding generation vector store management and index refresh cycles Implement versioning and change detection so that repositories remain synchronised with authoritative source systems Define and apply access controls at the data layer so that retrieval respects existing entitlement and confidentiality boundaries Measure and tune retrieval performance working with AI engineers to diagnose grounding failures and improve recall and precision
Master Data Readiness
Work with functional and technical teams to improve data quality and master data readiness across customer vendor product employee and asset domains Profile source data to quantify completeness consistency duplication and timeliness and report readiness objectively to initiative sponsors Implement automated data quality rules validation checks and exception reporting within pipelines Support remediation of root-cause data issues with business data owners rather than correcting symptoms downstream
Governance Trust Security
Ensure AI agents use trusted approved and governed data sources and that unapproved or unclassified content is excluded from AI consumption Apply data classification retention and privacy requirements in line with Group policy and UAE data protection regulation Maintain lineage and cataloguing so that any AI output can be traced back to its underlying source with confidence Collaborate with Cybersecurity and Compliance on access reviews data residency encryption and audit evidence
Platform Operations Collaboration
Operate and optimize cloud data platforms and pipelines for reliability performance and cost efficiency Monitor pipeline health resolve failures and maintain documentation runbooks and operational handover materials Partner with AI engineers application teams and business analysts throughout the delivery cycle from discovery to production support Contribute to Group data standards reusable pipeline patterns and shared engineering practice
Education
- – Bachelor’s degree in Computer Science, Information Systems, Data Engineering, Statistics, or a related discipline.
- – Postgraduate qualification in Data Science, Analytics, or Computer Science is an advantage.
Professional Certifications
- – Cloud data certification such as AWS Certified Data Engineer or Data Analytics, Microsoft Azure Data Engineer Associate, Google Cloud Professional Data Engineer, Databricks Data Engineer, or SnowPro.
- – Certification or formal training in data governance, data management (for example DAMA CDMP), or data privacy is advantageous.
- – Training in AI/GenAI data preparation, vector databases, or RAG architecture is an asset.
Experience
- – 5–8 years of experience in data engineering, data platforms, analytics, ETL/ELT, data warehousing, or cloud data solutions.
- – Hands-on experience in preparing structured and unstructured data for analytics, AI, GenAI, and Agentic AI use cases.
- – Experience with tools such as AWS S3, Glue, Redshift, Athena, Azure Data Factory, Synapse, Fabric, Databricks, Snowflake, BigQuery, or similar.
- – Demonstrated experience improving data quality and master data readiness in partnership with business functions.
- – Experience building and maintaining knowledge repositories or search/retrieval indexes is strongly preferred.
- – Exposure to multi-entity or group environments with heterogeneous source systems is an advantage.
Key Skills & Attributes
- – Strong SQL and Python capability with a disciplined, production-grade engineering approach.
- – Rigorous attention to data accuracy, lineage, and reproducibility.
- – Ability to assess and communicate data readiness honestly, including when a use case should not yet proceed.
- – Effective collaboration with business data owners to resolve issues at source.
- – Sound understanding of data privacy, classification, and security obligations.
- – Pragmatic balance between speed of delivery and long-term maintainability of data assets.