Senior Data Architect - Databricks / Data Lakehouse / GenAI
Senior Data Lakehouse Architect
Client: Bank Sector Client
About the Role
D L Resources is supporting a leading banking-sector client in hiring an experienced Senior Data Lakehouse Architect to lead the end-to-end architecture, design, and evolution of an enterprise Lakehouse platform.
The role will be responsible for defining the technical vision, target architecture, and engineering standards for modern data platforms, including data products, data marketplace, knowledge layers, real-time data processing, Generative AI, RAG, vector search, graph technologies, and agentic workloads.
The successful candidate should bring deep experience in large-scale enterprise data architecture, distributed computing, cloud and hybrid platforms, performance engineering, data governance, and modern DevSecOps practices.
Key Responsibilities
Own the end-to-end architecture and technical roadmap for the enterprise Data Lakehouse platform.
Design and evolve platform capabilities supporting:
Data products and data marketplace
Knowledge and semantic layers
Structured, semi-structured, and unstructured data
Real-time and streaming workloads
RAG and Generative AI use cases
Vector and graph-based data services
Agentic AI and autonomous workflow patterns
Define target architectures for applications and platform services with a focus on reusability, scalability, resilience, security, and operational efficiency.
Develop reusable architecture patterns, frameworks, and technical accelerators for:
Unstructured and multimodal content extraction
Batch and streaming architectures
Lambda and event-driven architectures
Retrieval-Augmented Generation (RAG)
Agentic workloads and AI-driven data processing
Partner with business and technology stakeholders to define data contracts, SLAs, data quality standards, and governance requirements for enterprise data products.
Provide architecture oversight and quality assurance to ensure solutions comply with the client’s software engineering, security, and delivery standards.
Review solution designs, technical specifications, non-functional requirements, and implementation approaches produced by engineering teams.
Participate in technology and product evaluations, proof-of-concepts, and RFP processes.
Guide installation, customization, integration, and operationalization of enterprise software platforms and technologies.
Lead performance engineering, capacity planning, scalability reviews, and optimization of distributed data workloads.
Partner with infrastructure, security, application, cloud, AI/ML, and operations teams to deliver integrated technology solutions.
Drive continuous service improvement, engineering automation, platform standardization, and operational excellence.
Produce architecture documentation, solution designs, implementation guidelines, operational standards, and technical runbooks.
Required Experience
10–15 years of experience in enterprise Data Engineering, Big Data, Data Architecture, Data Lake, or Lakehouse implementations.
Strong experience designing and delivering large-scale Data Lakehouse platforms, preferably within banking, financial services, or another highly regulated industry.
Proven experience across one or more leading data and cloud platforms such as:
Databricks, Snowflake, Cloudera, Azure, AWS, Google Cloud Platform, Huawei Cloud, or Alibaba Cloud.Strong experience designing distributed compute and MPP workloads across on-premise, hybrid, and cloud environments.
Deep understanding of enterprise data architecture, scalability, resilience, security, governance, and performance optimization.
Core Lakehouse & Data Architecture Skills
Strong experience in several of the following areas:
Open Table Formats: Apache Iceberg, Apache Hudi, Delta Lake
Object Storage: Cloud and enterprise object storage, including hot/warm/cold tiering strategies
Data Federation: Trino, Denodo, Dremio
Distributed Query Technologies: Hive, Impala, Apache Kudu and similar platforms
Data Processing: Spark, PySpark, SQL, Java, Python, Scala
Real-Time & Streaming: Apache Kafka, Confluent, Azure Event Hubs, Amazon Kinesis, Apache Flink, Spark Streaming, Structured Streaming, Apache NiFi
Workflow & Scheduling: Airflow, Control-M
Data Modelling & Governance: Enterprise data modelling, metadata, lineage, data contracts, data quality, and governance frameworks
Generative AI, RAG & Agentic Architecture
Experience designing or supporting modern AI-enabled data architectures, including:
Retrieval-Augmented Generation (RAG)
Embedding strategies and vectorization
Vector databases and vector search
Graph databases and knowledge graphs
Prompt and context management
Agentic workflow orchestration
Knowledge and semantic layers
AI-driven analytics and Generative BI
Relevant technologies may include:
Vector Search / Vector Databases
Databricks Vector Search
Azure AI Search
Pinecone
ChromaDB
Weaviate
Snowflake Cortex
Graph Databases
Neo4j
JanusGraph
TigerGraph
Microsoft Fabric / Cosmos DB
Amazon Neptune
Stardog
Agentic & AI Orchestration Frameworks
LangGraph
OpenAI Agents SDK
Microsoft Agent Framework
LlamaIndex Workflows
Google Agent Development Kit (ADK)
Data Products & Data Marketplace
Experience designing and delivering foundation and business data products.
Experience defining and implementing data contracts, service levels, governance, and quality controls.
Ability to expose data products through:
APIs
Publish/subscribe and event-driven architectures
Real-time dashboards
BI and Generative BI platforms
Data marketplace capabilities
Experience designing data products for enterprise consumption, reuse, discoverability, and governance.
Cloud & Hybrid Architecture
Strong understanding of cloud and hybrid architecture patterns, including:
Workload placement and cloud optimization strategies
Private and dedicated cloud connectivity such as AWS Direct Connect and Azure ExpressRoute
Data egress and network cost optimization
Infrastructure-as-Code
Hybrid and multi-cloud data architecture
Security and network integration
High availability and disaster recovery
DevOps, Platform Engineering & Automation
Experience with modern DevOps and software delivery practices, including:
CI/CD: Jenkins, Azure Pipelines, AWS CodePipeline, Google Cloud Build / Deploy
Source Control: Git, Bitbucket
Code Quality: SonarQube
Artifact Repositories: JFrog Artifactory, AWS CodeArtifact, Amazon ECR, Azure Artifacts, Google Artifact Registry
Infrastructure-as-Code: Terraform, AWS CloudFormation, Azure ARM
Containerization: Docker, Kubernetes, OpenShift
Deployment: Helm, Kustomize
Monitoring: AWS CloudWatch, Azure Monitor, Google Cloud Monitoring
Incident / Service Management: Remedy or equivalent platforms
Testing / Defect Management: JIRA, QuerySurge or similar tools
Programming & Automation
Strong knowledge of one or more of the following:
Python
Scala
Java
SQL
JavaScript / Node.js
Shell scripting
Groovy
Experience automating engineering and operational processes is strongly preferred.
Migration & Modernization Experience
Experience with migration and modernization programs involving legacy or MPP data platforms will be advantageous, including:
Teradata
Greenplum
Netezza
Other enterprise MPP platforms
Experience with bulk migration, workload modernization, automated migration tooling, and AI-assisted migration accelerators is a plus.
Education
Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
Equivalent relevant professional experience may also be considered.
Preferred Certifications
Candidates with relevant architecture, data, and cloud certifications will have an advantage. Examples include:
Databricks Certified Data Engineer / Data Architect
Microsoft Azure certifications
AWS Cloud / Data certifications
Google Cloud certifications
DAMA Certified Data Management Professional (CDMP)
Data modelling certifications such as Erwin
Relevant Kubernetes, DevOps, data engineering, or architecture certifications
What Will Help You Succeed
Strong architectural thinking with the ability to balance business outcomes, engineering quality, cost, scalability, security, and operational requirements.
Ability to understand enterprise-wide technology landscapes and translate them into practical technical roadmaps.
Strong analytical, troubleshooting, and decision-making capabilities.
Ability to resolve complex architecture and integration challenges.
Strong focus on engineering quality and continuous improvement.
Excellent communication skills, including the ability to explain complex technical concepts to non-technical stakeholders.
Strong stakeholder management and collaboration skills across business, technology, vendors, and distributed engineering teams.
Experience working in Agile and modern software delivery environments.
Ability to manage multiple initiatives and priorities in a fast-paced enterprise environment.
Key Technology Stack
Lakehouse & Data Platforms:
Databricks, Snowflake, Cloudera, Iceberg, Hudi, Delta Lake, Trino, Denodo, Dremio, Hive, Impala, Kudu
Cloud:
Azure, AWS, GCP, Huawei Cloud, Alibaba Cloud
Data Processing & Streaming:
Spark, PySpark, Python, Scala, Java, SQL, Kafka, Confluent, Flink, Spark Streaming, Structured Streaming, NiFi
AI / GenAI:
RAG, Vector Search, Embeddings, Graph Databases, Knowledge Graphs, LangGraph, OpenAI Agents SDK, LlamaIndex, Microsoft Agent Framework, Google ADK
DevOps & Platform Engineering:
Kubernetes, OpenShift, Docker, Terraform, Helm, Kustomize, Jenkins, Git, SonarQube, CI/CD
Data Architecture & Governance:
Data Products, Data Marketplace, Data Contracts, Data Quality, Metadata, Lineage, Data Modelling, Governance
Skills
- Agentic AI
- Agile
- AI
- Airflow
- Analytics
- API
- Artifactory
- Automation
- AWS
- Azure
- Bash
- Bitbucket
- ChromaDB
- CI/CD
- Cloud
- CloudFormation
- CloudWatch
- Containerization
- Data Engineering
- Data Governance
- Data Lake
- Data Modeling
- Data Quality
- Databricks
- Delta Lake
- DevOps
- DevSecOps
- Distributed Computing
- Docker
- Embeddings
- Event Driven Architecture
- Flink
- GCP
- Generative AI
- Git
- Groovy
- Helm
- Hive
- Iceberg
- Infrastructure as Code
- Java
- JavaScript
- Jenkins
- Jira
- Kafka
- Kinesis
- Kubernetes
- Lakehouse
- Lambda
- LangGraph
- LlamaIndex
- Machine Learning
- Microsoft Fabric
- Neo4j
- NiFi
- Node.js
- OpenAI
- OpenShift
- Pinecone
- PySpark
- Python
- RAG
- Scala
- Snowflake
- SonarQube
- Spark
- SQL
- Stakeholder Management
- Teradata
- Terraform
- Trino
- Vector Databases
- Vector Search
- Weaviate
- Workflow Orchestration
