Senior Data Engineer – PySpark, Databricks & Data Lakehouse
Senior Data Engineer / Lakehouse Engineer
to design, build, and operationalize enterprise-scale data platforms, Lakehouse solutions, data products, and data marketplace capabilities.
The successful candidate will have strong hands‑on expertise in modern data engineering technologies, distributed processing, cloud data platforms, streaming, open table formats, and DevSecOps practices. The role will also contribute to emerging
Generative AI, RAG, vector search, multimodal data processing, and AI-driven data pipelines .
Key Responsibilities
Design, implement, and operationalize enterprise
Data Lake/Lakehouse platforms, data products, and data marketplace capabilities .
Build scalable data ingestion and processing pipelines supporting
batch, streaming, CDC, event-driven, and API-based integration patterns .
Develop multimodal and unstructured data ingestion pipelines, including:
Content extraction from multiple document and file formats.
Metadata and field extraction using regular expressions and other processing techniques.
Extraction and processing of embedded images.
Video frame extraction and processing.
Audio and transcript extraction.
Build, test, deploy, and maintain foundation and business data products based on agreed
data contracts, SLAs, reconciliation rules, and data quality controls .
Implement and manage modern open table formats including
Apache Iceberg, Apache Hudi, and Delta Lake .
Develop data pipelines and platform capabilities supporting
RAG, vector search, Generative AI, NLP, and agentic AI use cases .
Optimize Spark and distributed data workloads for scalability, reliability, cost, and performance.
Perform production troubleshooting, performance tuning, root‑cause analysis, and operational support.
Implement data transformation, reconciliation, metadata, lineage, governance, and data quality frameworks.
Expose and distribute data through
APIs, event streams, dashboards, BI platforms, and data products .
Develop internal engineering tools and supporting applications using
Python, shell scripting, APIs, and modern web frameworks
where required.
Create and maintain technical architecture documentation, deployment guides, operating procedures, and production runbooks.
Ensure solutions comply with enterprise engineering standards, security requirements,
DevSecOps controls, CI/CD practices, and software delivery standards .
Collaborate with distributed engineering, architecture, analytics, AI/ML, business, and technology teams across multiple initiatives.
Required Experience & Technical Skills
8–12 years of experience
in Data Engineering, Big Data, Data Lake, Data Warehouse, or Lakehouse implementations.
Strong hands‑on experience with one or more enterprise data platforms such as: Databricks, Snowflake, Cloudera, Microsoft Azure, AWS, Google Cloud Platform (GCP), Huawei Cloud, or Alibaba Cloud .
Strong experience designing and developing
enterprise data products and/or data marketplace solutions .
Advanced hands‑on expertise in: Apache Spark, PySpark, SQL, Python, and/or Scala .
Strong programming skills in one or more of: Python, Scala, Java, and SQL .
Experience implementing open table formats such as
Apache Iceberg, Apache Hudi, and Delta Lake , together with cloud/object storage platforms.
Proven experience building enterprise frameworks for: data ingestion, transformation, reconciliation, validation, and data quality .
Hands‑on experience with relevant distributed data and integration technologies such as: Kafka, Flink, Spark Streaming, Airflow, Trino, Dremio, Hive, and Impala .
Strong experience with containerization, orchestration, infrastructure automation, and CI/CD technologies including: Kubernetes, OpenShift, Docker, Terraform, Jenkins, Git, and CI/CD pipelines .
Experience implementing monitoring, logging, observability, and production support capabilities for enterprise data platforms.
Strong understanding of
data modelling, metadata management, data lineage, governance, and data security .
Experience designing architectures for
structured, semi‑structured, and unstructured data
across Data Lake, Lakehouse, and Data Warehouse environments.
Experience exposing data through
REST APIs, event streams, dashboards, BI platforms, and other consumption channels .
AI / ML & Advanced Data Engineering Experience Experience in one or more of the following areas will be highly advantageous:
Building data architectures supporting
NLP, Generative AI, RAG, vector databases/vector search, and AI-driven analytics .
Ingestion, extraction, curation, enrichment, and governance of unstructured and multimodal data.
Experience with ML platforms and frameworks such as
MLflow, Cloudera Machine Learning (CML), Spark MLlib, scikit‑learn, and XGBoost .
Experience supporting machine‑learning model deployment and operationalization.
Building internal engineering applications or tools using
Python, Flask, React, shell scripting, or similar technologies .
Additional Advantageous Experience
Experience with enterprise migration or modernization involving
Teradata, Netezza, Greenplum, or other MPP data warehouse technologies .
Experience working within large‑scale banking, financial services, regulated, or enterprise environments.
Understanding of enterprise data governance, security, privacy, and regulatory requirements.
Education
Bachelor’s degree in
Computer Science, Engineering, Information Technology, Data Science , or a related discipline.
Preferred Certifications Relevant professional certifications are advantageous, including:
Databricks Certified Data Engineer
Microsoft Azure Data Engineer certification
AWS Data Engineering / Analytics certification
Google Cloud Professional Data Engineer
Snowflake SnowPro Certification
DAMA Certified Data Management Professional (CDMP)
What Will Help You Succeed
Strong
engineering, automation, and problem‑solving mindset .
Excellent troubleshooting, performance‑tuning, and root‑cause‑analysis capabilities.
Ability to design solutions for high‑volume, highly available, enterprise‑scale data environments.
Strong communication and stakeholder‑management skills.
Ability to collaborate effectively across distributed teams and manage multiple concurrent initiatives.
Experience working within
Agile and DevSecOps delivery environments .
Strong commitment to engineering quality, security, operational excellence, and continuous improvement.
Key Technology Stack Data Engineering & Processing: Spark, PySpark, Python, Scala, Java, SQL, Kafka, Flink, Spark Streaming, Airflow
Lakehouse & Data Platforms: Databricks, Snowflake, Cloudera, Iceberg, Hudi, Delta Lake, Trino, Dremio, Hive, Impala
Cloud & Platform Engineering: Azure, AWS, GCP, Huawei Cloud, Alibaba Cloud, Kubernetes, OpenShift, Docker, Terraform, Jenkins, Git, CI/CD
AI / ML & Advanced Analytics: RAG, Vector Search, Generative AI, NLP, MLflow, CML, Spark MLlib, scikit‑learn, XGBoost
Data Management: Data Products, Data Marketplace, Data Quality, Data Contracts, Metadata, Lineage, Governance, APIs, Event Streaming
Skills
- Agentic AI
- Agile
- AI
- Airflow
- Analytics
- API
- Automation
- AWS
- Azure
- Bash
- CI/CD
- Cloud
- Containerization
- Data Engineering
- Data Governance
- Data Ingestion
- Data Lake
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Quality
- Data Science
- Data Warehousing
- Databricks
- Delta Lake
- DevSecOps
- Docker
- Event Driven Architecture
- Flask
- Flink
- GCP
- Generative AI
- Git
- Hive
- Iceberg
- Java
- Jenkins
- Kafka
- Kubernetes
- Lakehouse
- Machine Learning
- Metadata Management
- MLflow
- Model Deployment
- NLP
- Observability
- OpenShift
- PySpark
- Python
- React
- REST
- Scala
- scikit-learn
- Snowflake
- Spark
- SQL
- Teradata
- Terraform
- Trino
- Vector Databases
- Vector Search
- XGBoost