Principal Databricks Lakehouse Architect
Summary
Design and lead enterprise-scale Databricks Lakehouse solutions, building scalable data pipelines and optimizing Spark workloads for analytics and AI on AWS/Azure/GCP.
We are seeking an experienced Senior Data Architect with deep expertise in Databricks, modern data platforms, and cloud technologies to architect and deliver enterprise-scale data and AI solutions. The ideal candidate will define end-to-end data architecture, lead large-scale data modernization initiatives, and drive best practices across data engineering, governance, security, and analytics.
Key Responsibilities
- Design and implement enterprise-scale data platforms using Databricks Lakehouse Architecture.
- Architect scalable batch and real-time data pipelines for high-volume data processing.
- Lead cloud data modernization and migration initiatives across AWS, Azure, or GCP.
- Define enterprise data models, architecture standards, and governance frameworks.
- Optimize Spark workloads for performance, scalability, and cost efficiency.
- Implement Delta Lake, Unity Catalog, and enterprise data governance best practices.
- Collaborate with business stakeholders, solution architects, and engineering teams to translate business requirements into technical solutions.
- Mentor technical teams and provide architectural guidance throughout the project lifecycle.
- Drive CI/CD, Infrastructure as Code, automation, monitoring, and platform reliability.
- Ensure compliance with enterprise security, data privacy, and regulatory requirements.
Mandatory Skills
- 10+ years of experience in Data Engineering, Data Architecture, or Big Data solutions.
- 5+ years of hands-on experience with Databricks.
- Strong expertise in Apache Spark (PySpark/Scala/Spark SQL).
- Extensive experience with Delta Lake, Unity Catalog, and Databricks Workflows.
- Experience designing enterprise Lakehouse/Data Warehouse architectures.
- Strong knowledge of ETL/ELT frameworks and data integration patterns.
- Hands-on experience with AWS, Azure, or Google Cloud Platform.
- Proficiency in SQL and Python (Scala is an added advantage).
- Experience with orchestration tools such as Airflow, Azure Data Factory, or similar.
- Strong understanding of data governance, security, metadata management, and data quality.
- Experience with DevOps, Git, CI/CD pipelines, and Infrastructure as Code (Terraform preferred).
- Excellent stakeholder management and solution design skills.
Preferred Qualifications
- Databricks Certified Professional Data Engineer or Databricks Certified Solution Architect.
- Cloud certifications (AWS, Azure, or GCP).
- Experience with streaming technologies such as Kafka or Structured Streaming.
- Exposure to AI/ML platforms and MLOps is an advantage.
- Experience in Banking, Financial Services, Insurance, or large enterprise environments is preferred.