Databricks Data Engineer
Key Responsibilities
- Design, build, and operate end-to-end data pipelines (ingestion → transformation → serving) on Databricks and/or Snowflake, following medallion/layered architecture patterns
- Configure and administer workspaces/accounts — compute policies, resource sizing, environment setup, and access hierarchy — per the SDP reference architecture
- Implement data governance controls: catalogue and schema design, RBAC, row/column-level security, data masking, and lineage tracking
- Set up CI/CD and infrastructure-as-code for pipeline deployment and environment promotion (dev → test → prod)
- Configure monitoring, telemetry, and audit logging to meet SDP's central observability and security posture requirements
- Support UAT, integration testing, and parallel-run validation during migration and go-live
- Produce handover documentation (runbooks, access lists, escalation procedures) for agency operations teams
- Work directly with client, agency stakeholders, and Principal (Databricks/Snowflake) solution architects throughout delivery
Required Technical Skills — Databricks
- Unity Catalog — catalogue/schema design, access control, and data lineage
- Lakeflow / Delta Live Tables for pipeline orchestration; Delta Lake table format
- Databricks SQL and cluster/workspace administration (compute policies, pools, cost management)
- Databricks Asset Bundles (DABs) and Databricks Repos for CI/CD
- PySpark / Spark SQL for large-scale data transformation
- Working knowledge of Databricks system tables (audit logs, billing/usage, query history) for observability
- Minimum 5 years of hands-on experience in Data Engineering, Data Platform Engineering, or related disciplines.
- Minimum 3 years of hands-on experience with Databricks involving data pipeline development, platform administration, governance, and optimization.
Required Technical Skills
- Strong SQL and Python (PySpark or general-purpose) for data engineering
- Data modeling — dimensional design, star/snowflake schemas, semantic layers
- CI/CD pipelines (e.g., Azure DevOps, GitHub Actions, GitLab CI) for data engineering workflows
- Infrastructure-as-code (Terraform preferred) for provisioning cloud data platform resources
- Hands-on experience on at least one hyperscaler — AWS, Azure, or Google Cloud
- Understanding of data security and compliance frameworks applicable to government/public-sector environments
Preferred Qualifications
- Databricks Certified Data Engineer Associate/Professional
- SnowPro Core, or SnowPro Advanced: Data Engineer
- Prior experience delivering on a government or regulated-sector data platform, or exposure to compliance frameworks such as IM8 is an add on
- Experience working as part of a System Integrator (SI) delivery team alongside a platform Principal is an add on