Data Engineer in Banking Domain
Summary
This Data Engineer role involves building and maintaining batch and streaming pipelines to ingest and standardize banking data for AI platforms. The position focuses on data quality, lineage, and security using the Microsoft Azure stack, Apache tools, and vector-store management.
Urgent requirement for Data Engineer in Banking Domain is required for our banking clients in Kuwait
Proven experience building batch and streaming pipelines on cloud platforms, integrating enterprise systems and unstructured content into analytics/AI platforms
- Hands-on with Microsoft Fabric, Synapse, Data Factory, Databricks, Purview, and Azure AI Search; plus Apache Airflow, Spark, Kafka, dbt, Great Expectations, PostgreSQL/pgvector, and open metadata/lineage tools.
- Skilled in classification, tagging, row/document-level security, encryption in transit/at rest, secrets management, masking/tokenisation, and full lineage/audit trails.
- Strong SQL and Python, with practical experience in data quality, lineage, metadata management, and troubleshooting data synchronization across systems.
Role Overview
This role builds the pipelines that ingest, clean, standardize and serve the bank's documents and data into the AI platform . Owns the vector-store feeds, data quality,lineage and the connectors that make internal systems consumable by the copilots, while implementing the access and classification controls that keep that data safein motion and at rest.
Role Experience
- 4+ years in data engineering building batch and streaming pipelines on cloud.
- Experience integrating enterprise/core systems and unstructured content (documents, SharePoint) into analytics or AI platforms and Microsoft cloud infrastructure.
- Strong SQL and Python; data-quality, lineage and metadata management in practice.
Core Skills & Capabilities
- Microsoft Stack (reference Build)
- Microsoft Fabric / Azure Synapse, Azure Data Factory, Azure Databricks, Azure Data Lake Storage.
- Microsoft Purview for data catalogue, classification and lineage; Azure AI Search indexing pipelines.
- Open-source / Custom Stack
- Apache Airflow / Spark / Kafka, dbt for transformation, Great Expectations for data quality.
- PostgreSQL / pgvector and open catalog/lineage tooling (OpenMetadata, DataHub).
- Security, Access & Data Management
- Implements data classification and tagging so downstream retrieval can enforce who-sees-what; propagates row- and document-level security.
- Encrypts data in transit and at rest, manages secrets/credentials in Key Vault or Vault, and never hard-codes access to source systems.
- Builds full data lineage and audit trails and applies masking/tokenization for PII and PCI data in non-production environments.
Skills: data,bank,data bricks