freehire launches on Product Hunt on 26 August.

Follow →

DE&A - Sr. Data Engineer - Databricks

We are looking for a hands-on Senior Data Engineer with 7 to 9 years of experience to design, build, and optimize large-scale data pipelines and lakehouse solutions on Databricks. The ideal candidate is a strong individual contributor who stays current with the latest Databricks platform capabilities — including Unity Catalog, Lakeflow (Delta Live Tables), Lakebase, and Databricks' expanding AI/agent tooling — and can apply them to solve real-world data engineering problems at scale.

Key Responsibilities:

  • Design, develop, and maintain scalable ETL/ELT pipelines on Databricks using PySpark, Spark SQL, and Delta Lake for batch and streaming workloads.

  • Build and manage declarative pipelines using Lakeflow / Delta Live Tables (DLT), including expectations, data quality checks, and change data capture (CDC).

  • Implement and manage data governance, access control, lineage, and data sharing using Unity Catalog across multiple workspaces and clouds.

  • Optimize Spark jobs and Databricks clusters for performance and cost, leveraging Photon, serverless compute, auto-scaling, and job clustering best practices.

  • Design and implement medallion architecture (bronze/silver/gold) data models and lakehouse patterns for analytics and ML consumption.

  • Work with Databricks Workflows (Jobs) to orchestrate multi-task pipelines, including dependency management, retries, and monitoring/alerting.

  • Integrate Databricks with cloud-native services (AWS/Azure/GCP) such as S3/ADLS/GCS, Kafka/Event Hubs/Kinesis, Glue/ADF, and IAM/Entra ID for secure, automated data flows.

  • Apply CI/CD practices for Databricks using Databricks Asset Bundles (DABs), Repos, and Git integration; automate deployments across dev/test/prod.

  • Evaluate and adopt newer Databricks capabilities — Lakebase (serverless Postgres on the lakehouse), Unity Catalog Metrics, Genie/Agent Bricks, Mosaic AI, and real-time/streaming enhancements — and recommend where they add value to existing pipelines.

  • Implement data quality, testing, and observability frameworks (e.g., Great Expectations, DLT expectations, Lakehouse Monitoring) to ensure trustworthy, production-grade data.

  • Collaborate with data scientists, analysts, and business stakeholders to understand requirements and translate them into robust, reusable data engineering solutions.

  • Mentor junior engineers, participate in code reviews, and contribute to engineering best practices, coding standards, and documentation.

  • Troubleshoot production data pipeline issues, perform root-cause analysis, and drive continuous improvement in reliability and performance.

Required Skills & Experience

  • 7–9 years of overall experience in Data Engineering, with at least 3–4 years of hands-on, production experience on the Databricks platform.

  • Strong programming skills in Python and/or Scala, with deep hands-on expertise in PySpark and Spark SQL.

  • Solid experience with Delta Lake (ACID transactions, time travel, schema evolution, optimize/vacuum/Z-ordering, liquid clustering).

  • Hands-on experience with Lakeflow / Delta Live Tables (DLT) for building declarative, quality-controlled pipelines.

  • Working knowledge of Unity Catalog for centralized governance, fine-grained access control, data lineage, and cross-workspace data sharing.

  • Experience with Databricks Workflows/Jobs for pipeline orchestration, scheduling, and monitoring.

  • Proficiency with at least one major cloud platform (AWS, Azure, or GCP) and its native storage/compute/security services.

  • Experience with streaming technologies such as Structured Streaming, Kafka, Event Hubs, or Kinesis.

  • Strong SQL skills, including performance tuning, partitioning strategies, and query optimization on large datasets.

  • Familiarity with CI/CD for data platforms — Databricks Asset Bundles, Git-based version control, Terraform, and automated testing/deployment pipelines.

  • Understanding of data modeling concepts (dimensional modeling, medallion/lakehouse architecture) and data warehousing fundamentals.

  • Demonstrated ability to stay current with the Databricks product roadmap (e.g., Unity Catalog enhancements, Lakebase, Genie/Agent Bricks, Mosaic AI, Lakehouse Monitoring, serverless compute) and apply relevant updates to existing systems.

  • Strong analytical, debugging, and performance-tuning skills across the Databricks/Spark stack.

  • Excellent communication skills with the ability to work directly with cross-functional stakeholders and, where applicable, mentor junior team members.

Good to Have

  • Databricks Certified Data Engineer Associate/Professional or Databricks Certified Associate/Professional Developer for Apache Spark certification.

  • Exposure to MLOps/MLflow, Mosaic AI, or Databricks' agent/GenAI tooling (Agent Bricks, Genie, AI/BI dashboards).

  • Experience with Lakebase or other Postgres/OLTP-on-lakehouse patterns for operational analytics use cases.

  • Experience with dbt, Airflow, or similar orchestration/transformation tools alongside Databricks.

  • Prior experience in a regulated or high-governance data environment (finance, healthcare, or similar).

  • Contributions to internal frameworks, reusable pipeline templates, or engineering best-practice documentation.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available