Data Engineer Lead
Summary
Lead a team to build and maintain ETL pipelines, Azure Data Platform, and data storage infrastructure using PySpark, Python, SQL, and Azure Databricks.
Add expected salary to your profile for insights
- Lead and oversee a mixed team of in-house and contracted data engineers to evolve and sustain current ETL pipelines, data storage infrastructure and Azure Data Platform. Scope of work covers:
- Liaise with internal business stakeholders and external suppliers to successfully deliver data engineering deliverables
- Partner with internal stakeholders to guarantee data engineering outputs meet defined quality standards, and enable smooth production rollout for multiple data products
- Allocate resources efficiently to carry out system implementation and ongoing maintenance activities
- Ensure comprehensive technical documentation is produced for all deliverables, including data dictionaries, data flows, pipeline architecture, data mapping rules and data asset inventories
- Conduct technical reviews and formally approve technical outputs prepared by team members and third-party vendors, including workload estimation, impact assessments, design documentation, technical specifications, plus SIT, deployment and system operational artefacts
- Provide backing to the Data Engineering Lead, and coordinate with other data capability teams to organise and prioritise ongoing workstreams
- Collaborate with Digital & IT teams and external suppliers to drive required system adjustments responding to upstream application updates and downstream business requirements
Key Requirement
- Bachelors Degree in Information Technology, Computer Science, Information Systems, Information Management or comparable technical discipline
- Minimum 5 years professional experience within data analytics platforms and data engineering, with proven track record in data warehousing, data modelling & design, data integration, data migration, ETL/ELT, BI and big data solution delivery
- Track record of leading small data engineering teams, together with vendor coordination and management experience
- Advanced proficiency in SQL and PySpark / Python is mandatory
- Hands-on experience processing structured and unstructured datasets using common big data formats (CSV, Parquet) and Hive Metastore; practical working knowledge of Azure Databricks is mandatory
- Practical experience deploying and managing cloud platforms built on Azure ecosystem (Data Lake Storage Gen2, Azure SQL Database, Data Factory, DevOps) is strongly preferred
- Extensive experience across the full Software Development Lifecycle 20; covering design, development, testing, rollout and formal documentation is strongly preferred
- Exposure to designing, building and maintaining controls for data security and data privacy is strongly preferred
- Experience supporting Machine Learning engineering workflows or Agentic AI deployments is strongly preferred
- Practical knowledge of data quality governance and benchmarking frameworks is strongly preferred
- Prior involvement in Databricks UC migration initiatives is advantageous
- Familiarity with Business Intelligence, data lineage and data catalog tooling is advantageous
- Working understanding of Agile / Scrum frameworks and project management platforms such as Jira is advantageous
- Confident presentation and communication capabilities; robust analytical thinking, problem-solving and stakeholder engagement skills
- Fluent written and verbal communication in English and Cantonese
Unlock job insights
Hirer responsiveness Salary match Number of applicants
Sales and Event Marketing Executive (Fresh Graduates/Early-career Talents)
MI Marketing