Data Lakehouse Architect
Key Responsibilities:
Please take note of the following must-have requirement:
“We need consultants who have implemented a large-scale Lakehouse on one of the following platforms (Cloudera, Huawei, Google (Borderless Lakehouse, Big Query, Looker), AWS (Outpost, EMR), Azure(Synapse)) and worked with Open table formats to implement a medallion architecture, implemented data-as-a-service using APIs/pub-sub/market place etc.”
- You will be responsible for the end-to-end architecture of the lake house platform.
This includes the design and implementation of data products, data marketplace, knowledge layer and enabling agentic workloads to run out of the lake house platform.
- You will also be responsible for quality assurance of the team’s delivery in conformance with the Bank-defined software delivery methodology and tools.
- You will partner with other technology functions to help deliver required technology solutions.
- Other responsibilities include:
Provide technical vision and create roadmaps for the lake house platform
Create the target architecture for an application / set of applications with emphasis on platforms, reusability, scalability and security
Create frameworks, technical features which helps in faster operationalisation of new patterns such as unstructured content extraction, lambda architecture deployment patterns, retrieval-augmented data patterns, agentic workloads etc
Effectively partner with business users to design data contracts, SLA, data quality rules for data products
Independently install, customise and integrate software packages and programs
Participate in selection of product/tools via RFP/POC.
Create technical documents (functional/non-functional specification, design specification, training manual) for the solutions. Review design specifications created by development team
- Performance engineering and tuning
Execute continuous service improvement and process improvement plans
Requirements:
EDUCATION
- Bachelor’s degree/University degree in Computer Science, Engineering, or equivalent experience
Certifications:
At least, 2 to 3 technical certifications in any of the below technologies:
1. Cloud/Lakehouse Certifications – (Azure, AWS, GCP cloud certifications)
2. DAMA Certified Data Management Professional, Databricks Certified Data Architect
3. Data modelling tools (Erwin)
4. Language– SQL, Java, Python, Scala, Javascript,node.js
5. Automation/ scripting – CtrlM, Shell Scripting, Groovy
6. VectorDB (Databricks Vector Search, Azure AI Search, Pinecone, ChromaDB, Weaviate, Snowflake Cortex)
7. GraphDB (Neo4J, Janusgraph, Tigergraph, Microsoft Fabric + Cosmos DB, Amazon Neptune, Stardog)
8. Agentic Orchestration/Harness and Agentic Flow Frameworks (LangGraph, OpenAI AgentsSDK, Microsoft Agent Framework, LlamaIndex Workflows, Google ADK)
9. NOSQL/ In-memory DB (Azure Cosmos DB, Redis, Amazon DynamoDB, Firestore, Redis/Valkey)
10. Event Streaming (Apache Kafka, Confluent, Event Hubs, Kinesis)
11. Real-time processing (Flink, Spark Streaming, NiFi, Structured Streaming)
Additional Experience required for all teams to create an added advantage:
1. CI/CD software - Jenkins, JIRA, Code pipelines, Azure pipelines, GCP Cloud Build +Deploy
2. Code Quality – SonarQube
3. Artifact Repository – Jfrog, Code Artefact, ECR, Azure Artifacts, GCP Artifact Registry
4. Source Code Version Control Tool – Git, Bitbucket
5. Infrastructure-as-code– Terraform, Cloud formation, ARM
6. Deployment Tool kit -Jenkins
7. Monitoring– CloudWatch, Azure Monitor, Cloud Monitoring
8. Service or Incident Management (IcM) Tools - Remedy
9. Scheduling Tool - Control-M, Airflow
10. Defect Management Tool - JIRA
11. Application Testing tool – Query Surge
Key Domain/ Technical Skills:
TEAM Architecture(Big Data)
- 10-15years of experience of implementing a Data Lakehouse preferably in FSI domain (using platforms such as Databricks, Snowflake, Cloudera, Huawei, Alibaba, Google Cloud, AWS, Azure),
- Experience in large scale implementations and performance optimizations in the Lakehouse using
a. OpenTable Formats such as Iceberg, Hudi, Delta Lake,
b. Object Storage including tiered storage (hot, warm, cold data) strategies
c. Data Federation such as Trino, Denodo, Dremio
d. Multimodal Query Engines (Hive, Impala, Apache Kudu etc)
- Experience in designing MPP and Distributed Compute workloads across on-premise, hybrid and cloud environments
- Experience in serving agentic workloads using RAG, Embedding strategies, Vector DB, GraphDB, prompt engineering, context management etc
- Experience in designing optimal hybrid and cloud workloads using private dedicated connectivity (Direct Connect, Express Route etc), workload placement strategy, egress cost optimization, Infrastructure-as-Code,
- Experience in building foundation and business data products and serving them to downstream applications via API, pub-and-sub, generative BI, real-time dashboards, etc and publishing to a data marketplace
- Knowledge of migrating workloads out of MPP appliances such as Teradata, Greenplum, Netezza using bulk migration strategies, agentic accelerators is a plus
- Knowledge of containerization, deploying applications to Kubernetes, Openshift using Helm package manager, Kustomize etc is a plus
Expertise in integrating applications with Devops tools
Skills
- Agentic AI
- AI
- Airflow
- API
- Automation
- AWS
- Azure
- Bash
- Bitbucket
- ChromaDB
- CI/CD
- Cloud
- CloudWatch
- Containerization
- Data Modeling
- Data Quality
- Databricks
- Delta Lake
- DevOps
- DynamoDB
- Express.js
- Flink
- GCP
- Git
- Groovy
- Helm
- Hive
- Infrastructure as Code
- Java
- JavaScript
- Jenkins
- Jira
- Kafka
- Kinesis
- Kubernetes
- Lakehouse
- Lambda
- LangGraph
- LlamaIndex
- Looker
- Microsoft Fabric
- Neo4j
- NiFi
- Node.js
- NoSQL
- OpenAI
- OpenShift
- Pinecone
- Process Improvement
- Prompt Engineering
- Python
- Redis
- Scala
- Snowflake
- SonarQube
- Spark
- SQL
- Teradata
- Terraform
- Trino
- Vector Search
- Version Control
- Weaviate