DE&A - Sr. Data Engineer - Databricks
We are looking for a hands-on Senior Data Engineer with 7 to 9 years of experience to design, build, and optimize large-scale data pipelines and lakehouse solutions on Databricks. The ideal candidate is a strong individual contributor who stays current with the latest Databricks platform capabilities — including Unity Catalog, Lakeflow (Delta Live Tables), Lakebase, and Databricks' expanding AI/agent tooling — and can apply them to solve real-world data engineering problems at scale.
Key Responsibilities:
Design, develop, and maintain scalable ETL/ELT pipelines on Databricks using PySpark, Spark SQL, and Delta Lake for batch and streaming workloads.
Build and manage declarative pipelines using Lakeflow / Delta Live Tables (DLT), including expectations, data quality checks, and change data capture (CDC).
Implement and manage data governance, access control, lineage, and data sharing using Unity Catalog across multiple workspaces and clouds.
Optimize Spark jobs and Databricks clusters for performance and cost, leveraging Photon, serverless compute, auto-scaling, and job clustering best practices.
Design and implement medallion architecture (bronze/silver/gold) data models and lakehouse patterns for analytics and ML consumption.
Work with Databricks Workflows (Jobs) to orchestrate multi-task pipelines, including dependency management, retries, and monitoring/alerting.
Integrate Databricks with cloud-native services (AWS/Azure/GCP) such as S3/ADLS/GCS, Kafka/Event Hubs/Kinesis, Glue/ADF, and IAM/Entra ID for secure, automated data flows.
Apply CI/CD practices for Databricks using Databricks Asset Bundles (DABs), Repos, and Git integration; automate deployments across dev/test/prod.
Evaluate and adopt newer Databricks capabilities — Lakebase (serverless Postgres on the lakehouse), Unity Catalog Metrics, Genie/Agent Bricks, Mosaic AI, and real-time/streaming enhancements — and recommend where they add value to existing pipelines.
Implement data quality, testing, and observability frameworks (e.g., Great Expectations, DLT expectations, Lakehouse Monitoring) to ensure trustworthy, production-grade data.
Collaborate with data scientists, analysts, and business stakeholders to understand requirements and translate them into robust, reusable data engineering solutions.
Mentor junior engineers, participate in code reviews, and contribute to engineering best practices, coding standards, and documentation.
Troubleshoot production data pipeline issues, perform root-cause analysis, and drive continuous improvement in reliability and performance.
Required Skills & Experience
7–9 years of overall experience in Data Engineering, with at least 3–4 years of hands-on, production experience on the Databricks platform.
Strong programming skills in Python and/or Scala, with deep hands-on expertise in PySpark and Spark SQL.
Solid experience with Delta Lake (ACID transactions, time travel, schema evolution, optimize/vacuum/Z-ordering, liquid clustering).
Hands-on experience with Lakeflow / Delta Live Tables (DLT) for building declarative, quality-controlled pipelines.
Working knowledge of Unity Catalog for centralized governance, fine-grained access control, data lineage, and cross-workspace data sharing.
Experience with Databricks Workflows/Jobs for pipeline orchestration, scheduling, and monitoring.
Proficiency with at least one major cloud platform (AWS, Azure, or GCP) and its native storage/compute/security services.
Experience with streaming technologies such as Structured Streaming, Kafka, Event Hubs, or Kinesis.
Strong SQL skills, including performance tuning, partitioning strategies, and query optimization on large datasets.
Familiarity with CI/CD for data platforms — Databricks Asset Bundles, Git-based version control, Terraform, and automated testing/deployment pipelines.
Understanding of data modeling concepts (dimensional modeling, medallion/lakehouse architecture) and data warehousing fundamentals.
Demonstrated ability to stay current with the Databricks product roadmap (e.g., Unity Catalog enhancements, Lakebase, Genie/Agent Bricks, Mosaic AI, Lakehouse Monitoring, serverless compute) and apply relevant updates to existing systems.
Strong analytical, debugging, and performance-tuning skills across the Databricks/Spark stack.
Excellent communication skills with the ability to work directly with cross-functional stakeholders and, where applicable, mentor junior team members.
Good to Have
Databricks Certified Data Engineer Associate/Professional or Databricks Certified Associate/Professional Developer for Apache Spark certification.
Exposure to MLOps/MLflow, Mosaic AI, or Databricks' agent/GenAI tooling (Agent Bricks, Genie, AI/BI dashboards).
Experience with Lakebase or other Postgres/OLTP-on-lakehouse patterns for operational analytics use cases.
Experience with dbt, Airflow, or similar orchestration/transformation tools alongside Databricks.
Prior experience in a regulated or high-governance data environment (finance, healthcare, or similar).
Contributions to internal frameworks, reusable pipeline templates, or engineering best-practice documentation.