Senior Data Engineer
Summary
Senior Data Engineer in Kuala Lumpur who designs, develops, tests, and deploys ETL and data consolidation solutions, optimizes SQL queries and data models, and provides BAU/production support for data platforms. Core focus on ETL, SQL, and data modelling, posted via a recruitment consultancy.
Job Responsibilities:
- Design, develop, and maintain scalable data pipelines and data solutions using Databricks and PySpark.
- Build and optimize ETL/ELT solutions for batch and near real-time data processing.
- Design and maintain enterprise Data Warehouse and Lakehouse solutions, including Delta Lake and Delta Tables.
- Develop dimensional data models using Fact/Dimension tables, Star Schema, Snowflake Schema, and SCD Type 1 & 2.
- Implement and maintain data governance and security using Unity Catalog, RBAC, ABAC, data lineage, auditing, and fine-grained access controls.
- Configure and manage Delta Sharing for secure internal and external data collaboration.
- Support Databricks Genie and Genie Spaces adoption and administration.
- Perform PySpark, SQL, Data Warehouse and Lakehouse performance tuning, including query, cluster and workload optimization.
- Develop and maintain CI/CD pipelines and automate deployment, testing and release processes using GitHub, GitHub Actions or equivalent tools.
- Collaborate with business users, data analysts, architects and technology teams to translate business requirements into scalable data solutions.
- Provide technical guidance and mentorship to junior team members where required.
- Continuously improve data platform performance, scalability, security, governance and cost efficiency.
Job Requirements:
- Bachelor's Degree in Computer Science, Information Technology, Engineering, Data Science or a related field.
- 5 years and above of experience in Data Engineering, Data Warehousing or Big Data technologies.
- Strong hands-on experience with Databricks, Azure Databricks, PySpark, Delta Lake, Delta Tables and Databricks Workflows.
- Strong knowledge of Unity Catalog, RBAC, ABAC, Delta Sharing, data governance and data security.
- Experience with Databricks Genie / Genie Spaces is preferred.
- Strong SQL skills with experience in query and data warehouse performance optimization.
- Solid experience in Data Warehouse architecture, Lakehouse architecture and Medallion design patterns.
- Strong knowledge of dimensional data modelling, including Fact/Dimension tables, Star Schema, Snowflake Schema and SCD.
- Experience with GitHub, Git-based workflows and CI/CD, preferably using GitHub Actions.
- Experience with at least one major cloud platform (Azure, AWS or GCP).
- Knowledge of Infrastructure-as-Code, preferably Terraform, data observability and monitoring is an advantage.
- Databricks certification is an added advantage.
- Strong analytical, problem-solving and troubleshooting skills.
- Good communication and stakeholder management skills, with the ability to work independently and across business and technical teams.
Skills
- AWS
- Azure
- CI/CD
- Cloud
- Data Engineering
- Data Governance
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Science
- Data Warehousing
- Databricks
- Delta Lake
- Design Patterns
- ELT
- ETL
- GCP
- Git
- GitHub
- GitHub Actions
- Infrastructure as Code
- Lakehouse
- Observability
- PySpark
- RBAC
- Snowflake
- SQL
- Stakeholder Management
- Terraform
- Unity