Lead Enterprise Lakehouse Architect – Data Products & Agentic AI- Contract
NTT SINGAPORE PTE. LTD. Lead Enterprise Lakehouse Architect – Data Products & Agentic AI- Contract
Lead Enterprise Lakehouse Architect – Open Table Formats, Data Products & Agentic AI
Contract Duration: 09 months
Seniority: L4 – More than 10 years of relevant experience
Working Arrangement: Onsite ( 5 days from office )
Headcount: 1
Role Overview
We are seeking an experienced Enterprise Data Lakehouse Architect to own the end-to-end architecture of a large-scale Lakehouse platform supporting governed data products, Data-as-a-Service, real-time analytics, knowledge layers and agentic AI workloads.
This is a senior hands-on architecture position requiring demonstrable production implementation experience. Applicants whose experience is limited to traditional data warehouses, BI reporting, general cloud architecture or data-engineering delivery without end-to-end Lakehouse ownership will not meet the requirements.
Responsibilities
- Define the technical vision, target architecture and implementation roadmap for an enterprise-scale Lakehouse platform.
- Architect reusable, scalable and secure platform components across on-premises, hybrid and cloud environments.
- Design and implement Bronze, Silver and Gold medallion layers using Delta Lake, Apache Iceberg or Apache Hudi.
- Design object-storage architecture covering lifecycle management and hot, warm and cold data-tiering strategies.
- Architect MPP and distributed-compute workloads using Spark, Databricks, BigQuery, Dataproc, EMR, Synapse or equivalent platforms.
- Establish foundation and business data products with formal data contracts, SLAs, ownership, lineage and data-quality rules.
- Serve governed data products to downstream applications through REST APIs, Kafka/Pub-Sub, real-time streams, dashboards and data-marketplace capabilities.
- Design reusable patterns for structured and unstructured content ingestion, lambda processing and retrieval-augmented data workloads.
- Enable RAG and agentic AI workloads using embeddings, vector databases, graph databases, prompt engineering and context-management strategies.
- Design secure hybrid-cloud connectivity using private dedicated connectivity, workload-placement strategies and data-egress cost controls.
- Implement Infrastructure-as-Code and automated platform provisioning.
- Lead platform performance engineering, query optimisation, capacity planning, reliability improvements and FinOps initiatives.
- Evaluate Lakehouse, federation, query-engine, vector-database and graph-database technologies through RFPs and proofs of concept.
- Define functional, non-functional, security and solution-design specifications.
- Review technical designs and delivery outputs for compliance with architecture, engineering, security and quality standards.
- Integrate the Lakehouse platform with enterprise CI/CD, testing, source-control, monitoring, scheduling and incident-management tools.
- Lead continuous service-improvement and process-improvement initiatives.
Mandatory Requirements
Applicants must meet all the following requirements:
- Between 10 and 15 years of relevant experience in enterprise data architecture, big-data platforms and distributed data processing.
- At least five years of hands-on architecture ownership for enterprise-scale data platforms.
- Personally architected and implemented at least one production-scale Lakehouse in banking or financial services.
- Hands-on implementation experience with at least one approved platform:ClouderaHuawei CloudGoogle BigQuery, BigLake, Dataplex or DataprocAWS EMR or OutpostsAzure Synapse or Azure Databricks
- Production implementation of Bronze, Silver and Gold medallion architecture.
- Deep hands-on experience with at least one open-table format: Delta Lake, Apache Iceberg or Apache Hudi.
- Ability to explain ACID transactions, schema evolution, partition evolution, time travel/snapshots, compaction and small-file management.
- Experience designing distributed Spark/PySpark workloads and performing query, storage and compute optimisation.
- Production experience implementing both batch and real-time/streaming pipelines.
- Hands-on Data-as-a-Service implementation using REST APIs and Kafka/Pub-Sub.
- Experience building reusable foundation and business data products supported by data contracts, SLAs and automated data-quality controls.
- Experience publishing governed data products through a catalogue, exchange or data marketplace.
- Experience with enterprise object storage and hot, warm and cold lifecycle strategies.
- Experience implementing metadata management, data lineage, RBAC, audit logging and fine-grained access controls.
- Production experience enabling RAG workloads using embeddings and a vector database.
- Practical knowledge of graph databases, prompt engineering, context management and LLM governance.
- Experience designing hybrid-cloud platforms, private connectivity, workload placement and egress-cost optimisation.
- Hands-on Infrastructure-as-Code experience using Terraform, CloudFormation or ARM/Bicep.
- Strong CI/CD implementation experience using Jenkins, Azure DevOps, Cloud Build, GitHub Actions or equivalent.
- Experience with platform monitoring, incident management, performance engineering and continuous service improvement.
- Ability to work onsite at IH2, Malaysia throughout the 12-month assignment.
Mandatory Certifications
Applicants must possess at least two current professional certifications, including:
- One professional-level cloud architecture or data-engineering certification from Google Cloud, AWS or Microsoft Azure; and
- One Databricks Data Engineer Professional, Databricks Data Architect, CDMP or equivalent data-platform certification.
Associate-level training badges or course-completion certificates alone will not satisfy this requirement.
Preferred Experience
- Trino, Denodo or Dremio data federation.
- Hive, Impala or Apache Kudu query engines.
- Migration from Teradata, Greenplum or Netezza into a modern Lakehouse.
- Databricks Vector Search, Azure AI Search, Pinecone, Weaviate, ChromaDB or Snowflake Cortex.
- Neo4j, JanusGraph, TigerGraph, Amazon Neptune or Stardog.
- LangGraph, OpenAI Agents SDK, Microsoft Agent Framework, LlamaIndex Workflows or Google ADK.
- Kubernetes or OpenShift deployment using Helm or Kustomize.
- Banking regulatory requirements and controls covering MAS, BCBS 239, AML, data residency and auditability.
Interested candidates are kindly requested to email their CV with their experience to sandeep.sringeripai@global.ntt
We look forward to your application!
Skills
- Agentic AI
- AI
- Analytics
- API
- AWS
- Azure
- Azure DevOps
- Bicep
- BigQuery
- ChromaDB
- CI/CD
- Cloud
- CloudFormation
- Data Engineering
- Data Lineage
- Data Quality
- Databricks
- Delta Lake
- DevOps
- Embeddings
- FinOps
- GCP
- GitHub
- GitHub Actions
- Helm
- Hive
- Iceberg
- Infrastructure as Code
- Jenkins
- Kafka
- Kubernetes
- Lakehouse
- Lambda
- LangGraph
- LlamaIndex
- LLM
- Metadata Management
- Neo4j
- OpenAI
- OpenShift
- Pinecone
- Process Improvement
- Prompt Engineering
- PySpark
- RBAC
- REST
- Snowflake
- Solution Design
- Spark
- Teradata
- Terraform
- Trino
- Vector Databases
- Vector Search
- Weaviate