freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer

Summary

Build and maintain scalable ETL/ELT pipelines on Azure using Databricks, PySpark, and Azure Data Factory to ingest, transform, and deliver clean data for analytics and ML.

Responsibilities

  • Design, develop, test, deploy and maintain ETL/ELT pipelines using Azure Data Factory (ADF) and Databricks.
  • Implement scalable data processing and transformation logic in PySpark and SQL on Databricks (Delta Lake / Lakehouse patterns).
  • Ingest and integrate data from diverse batch and streaming sources into ADLS Gen2 and Delta Lake.
  • Build curated, documented datasets for analytics and ML (support dimensional models where appropriate).
  • Implement automated data quality checks, unit/integration tests, and validation gates.
  • Monitor pipeline health, implement logging, alerting and dashboards; perform RCA and production incident remediation.
  • Optimise pipeline performance and cost cluster sizing, job scheduling, partitioning, caching and query tuning.
  • Implement and maintain CI/CD for data jobs and infrastructure (Terraform/ARM/Bicep, Azure DevOps or GitHub Actions).
  • Ensure data governance, lineage and security best practices; register assets in a data catalogue (e.g., Microsoft Purview or equivalent).
  • Collaborate with cross-functional teams to collect requirements, define SLAs/SLOs and deliver production-ready solutions.
  • Contribute to runbooks, documentation, coding standards and data platform roadmap.
  • Additional responsibilities for Senior Data Engineers: Lead design reviews, define platform standards and approve infra decisions; drive cost-control initiatives and SLA reporting; participate in vendor selection, proof-of-concepts and platform architecture decisions.

Qualifications

  • Bachelor’s degree in Information Systems, Computer Science, Data Management, or related field (or equivalent experience).
  • 3+ or 5+ years, with demonstrable leadership in delivery and design.
  • Hands-on Databricks experience (PySpark, SQL) and familiarity with Databricks runtimes.
  • Strong experience with Azure data services: Azure Data Factory, ADLS Gen2.
  • Strong SQL skills and production Python experience.
  • Experience with CI/CD
  • Experience applying data quality testing, monitoring and observability in production.
  • Strong debugging, performance tuning, and troubleshooting skills.
  • Good communication skills; experience working in cross-functional teams.
  • Fluent English and Mandarin.
  • Experience with Delta Lake, Lakehouse architectures and ACID semantics.
  • Familiarity with data cataloguing/governance tools (Microsoft Purview, Alation).
  • Certifications: Databricks Certified Data Engineer, Microsoft DP-203 or Azure certifications.
  • Experience with BI tools (Power BI, Looker, Tableau) and consuming datasets for analytics.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available