freehire launches on Product Hunt on 26 August.

Follow →

Senior Databricks Data Engineer

Summary

Senior hands-on role building and optimizing PySpark/SQL pipelines on Databricks, tuning Spark performance, and enforcing data governance in a Lakehouse architecture.

Brampton East, Canada | Posted on 08/13/2026

We are looking for a Senior, Super Hands-On DatabricksData Engineer who lives and breathes code, query optimization, and moderndata architecture. In this role, you won't just design architectures onwhiteboards—you will write production PySpark/SQL, optimize Databricksclusters, build streaming and batch pipelines, and enforce data governance.

You will own end-to-end pipeline execution from rawingestion to curated Gold layer models, playing a lead role in modernizing ourLakehouse platform.

Key Responsibilities

1. Hands-On Pipeline Development & LakehouseArchitecture

  • Design,build, and maintain enterprise-scale batch and real-time streamingpipelines using PySpark, SQL, Delta Live Tables (DLT), and AutoLoader.
  • Implementand refine Medallion Architecture (Bronze Silver Gold) to support downstream BI,reporting, and Machine Learning workloads.
  • Enforceschema evolution, ACID transactions, and data compaction using DeltaLake core constructs.

2. Performance Tuning & Optimization (Deep Tech)

  • Diagnoseand resolve Spark performance bottlenecks: data skew, OOM errors,excessive shufflings, and memory spills.
  • Optimizequeries using Liquid Clustering, Z-Ordering, Data Partitioning, AQE(Adaptive Query Execution), and Photon engine tuning.
  • Benchmarkand optimize Databricks compute workloads to minimize DBU (DatabricksUnit) consumption and cloud costs (FinOps).

3. Governance, Security & Quality

  • Implementend-to-end data governance, fine-grained access control (row/column-levelsecurity), and lineage tracking using Unity Catalog.
  • Automateautomated data quality validation checks and alert mechanisms across thepipeline life cycle.

4. Operations, CI/CD & DevOps

  • Automatepipeline orchestration using Databricks Asset Bundles (DABs) or DatabricksWorkflows / Apache Airflow.
  • BuildCI/CD pipelines (GitHub Actions, Azure DevOps, or GitLab) for automatedtesting, deployment, and code promotions.

Requirements

Required Skills &Qualifications

Must-Haves

  • Experience:8+ years in Data Engineering, with 4+ years of intensive, hands‑onproduction experience on Databricks.
  • ProgrammingMastery: Fluent in PySpark, Advanced SQL, and Python.
  • DatabricksEcosystem: Deep experience with Delta Lake, Unity Catalog, DeltaLive Tables (DLT), Auto Loader, and Databricks Workflows.
  • CloudInfrastructure: Strong hands‑on experience in at least one primarycloud provider (AWS, Azure, or GCP) integration with Databricks(S3/ADLS Gen2, IAM, Key Vaults/Secret Manager).
  • DataModeling: Solid understanding of dimensional modeling (Kimball), OneBig Table (OBT) strategies, and data vault patterns.
  • CI/CD& Software Engineering: Proficient in Git workflows, unit testingPySpark code (pytest), and deployment automation.

Preferred / Nice-to-Haves

  • Certifications: Databricks Certified Data Engineer Professional.
  • Streaming: Hands‑on with Apache Kafka, Event Hubs, or Kinesis integration viaStructured Streaming.
  • GenAI/ ML Ops: Familiarity with MLflow, Feature Store, or Vector Searchwithin Databricks.
  • Infrastructureas Code (IaC): Experience using Terraform to provision Databricksworkspaces and storage resources.

Performance Indicators(How success is measured)

  • PipelineReliability: Maintaining strict SLA thresholds on critical Gold-layermodels.
  • CostEfficiency: Measurable reduction in DBU costs through effectivecompute profiling and tuning.
  • CodeQuality: High test coverage and zero-downtime CI/CD deployments.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available