Data Engineer
Summary
Build and maintain scalable ETL/ELT pipelines on Azure using Databricks, PySpark, and Azure Data Factory to ingest, transform, and deliver clean data for analytics and ML.
Responsibilities
- Design, develop, test, deploy and maintain ETL/ELT pipelines using Azure Data Factory (ADF) and Databricks.
- Implement scalable data processing and transformation logic in PySpark and SQL on Databricks (Delta Lake / Lakehouse patterns).
- Ingest and integrate data from diverse batch and streaming sources into ADLS Gen2 and Delta Lake.
- Build curated, documented datasets for analytics and ML (support dimensional models where appropriate).
- Implement automated data quality checks, unit/integration tests, and validation gates.
- Monitor pipeline health, implement logging, alerting and dashboards; perform RCA and production incident remediation.
- Optimise pipeline performance and cost cluster sizing, job scheduling, partitioning, caching and query tuning.
- Implement and maintain CI/CD for data jobs and infrastructure (Terraform/ARM/Bicep, Azure DevOps or GitHub Actions).
- Ensure data governance, lineage and security best practices; register assets in a data catalogue (e.g., Microsoft Purview or equivalent).
- Collaborate with cross-functional teams to collect requirements, define SLAs/SLOs and deliver production-ready solutions.
- Contribute to runbooks, documentation, coding standards and data platform roadmap.
- Additional responsibilities for Senior Data Engineers: Lead design reviews, define platform standards and approve infra decisions; drive cost-control initiatives and SLA reporting; participate in vendor selection, proof-of-concepts and platform architecture decisions.
Qualifications
- Bachelor’s degree in Information Systems, Computer Science, Data Management, or related field (or equivalent experience).
- 3+ or 5+ years, with demonstrable leadership in delivery and design.
- Hands-on Databricks experience (PySpark, SQL) and familiarity with Databricks runtimes.
- Strong experience with Azure data services: Azure Data Factory, ADLS Gen2.
- Strong SQL skills and production Python experience.
- Experience with CI/CD
- Experience applying data quality testing, monitoring and observability in production.
- Strong debugging, performance tuning, and troubleshooting skills.
- Good communication skills; experience working in cross-functional teams.
- Fluent English and Mandarin.
- Experience with Delta Lake, Lakehouse architectures and ACID semantics.
- Familiarity with data cataloguing/governance tools (Microsoft Purview, Alation).
- Certifications: Databricks Certified Data Engineer, Microsoft DP-203 or Azure certifications.
- Experience with BI tools (Power BI, Looker, Tableau) and consuming datasets for analytics.