Data Engineer
Summary
Data engineering lead responsible for project delivery and technical architecture within a health business unit. Key technologies include Azure Databricks, Spark, and AliCloud big data tools (DataWorks, MaxCompute) for building scalable ETL/ELT pipelines and data lakehouse architectures.
Must-have criteria:
- Minimum 8+ years of IT experience, with at least 5+ years specifically in data engineering delivery leadership or technical lead roles - with the ability to independently own and drive project delivery in a timely manner
- Agile methodology
- Hands-on experience with Azure Cloud Services, database technology (Oracle, SQL Server, or PostgreSQL), Azure Databricks, Apache Spark (PySpark / Spark SQL), and modern ETL/ELT pipeline design, AliCloud big data tools, specifically DataWorks, MaxCompute, OSS, and Hologres
- Must be proficient in written and spoken English – only local candidates who do not require visa sponsorship will be considered
Good-to-have criteria:
- Experience with FineBI, Power BI, Tableau or similar enterprise reporting tools
- Any other language proficiencies
Position Responsibilities:
- Partner closely with Health business unit leaders, Health product owners, and IT stakeholders to capture business requirements, define project scopes, and establish clear delivery timelines.
- Partner and collaborate with Health Regional teams.
- Perform workload and effort estimation; break down complex data initiatives into detailed technical epics, user stories, and execution roadmaps.
- Design, build, and optimize scalable batch and real-time data pipelines, data lakehouse architectures, and data warehousing solutions.
- Provide technical leadership for enterprise data platforms, driving architecture, engineering best practices, platform optimization, and innovation across Azure Cloud and Databricks (Spark)
- Implement robust data orchestration, data ingestion, cleansing, transformation, augmentation, and data quality control processes.
- Establish and enforce technical quality standards, data governance frameworks, and best practices across data ingestion, storage, and processing.
- Conduct thorough code reviews, lead technical troubleshooting, and ensure data pipelines are secure, resilient, and cost-optimized