Senior Data Engineer
NewBe an early applicantSummary
Senior Data Engineer building and maintaining scalable data pipelines, ETL/ELT, and enterprise Data Warehouse/Lakehouse platforms on Databricks. Day-to-day spans PySpark and SQL development, Delta Lake modeling, Unity Catalog governance, performance tuning, CI/CD with GitHub Actions, and production support.
- Design, develop, and maintain scalable data pipelines and data products using Databricks and PySpark.
- Build and optimize ETL/ELT solutions supporting batch and near real-time data processing.
- Design, develop, and maintain enterprise Data Warehouse and Lakehouse solutions that support reporting, analytics, and AI/ML use cases.
- Develop and maintain enterprise data models using Delta Lake and Delta Tables.
- Design and implement dimensional data models, including Fact and Dimension tables, Star Schema, Snowflake Schema, and Slowly Changing Dimensions (SCD).
- Ensure data quality, reliability, scalability, and performance across the data platform.
- Implement best practices for code management, testing, deployment, and operational monitoring.
Databricks Platform & Governance
- Implement and manage Unity Catalog for centralized governance, data discovery, and security.
- Design and maintain governance frameworks utilizing:
- Role-Based Access Control (RBAC)
- Attribute-Based Access Control (ABAC)
- Fine-grained data permissions
- Data lineage and auditing
- Configure and manage Delta Sharing to support secure external and internal data collaboration.
- Support adoption and administration of Databricks Genie, including:
- Security and access governance controls
Performance Optimization
- Perform advanced PySpark performance tuning and troubleshooting.
- Optimize query performance, cluster utilization, partitioning strategies, and workload management.
- Identify bottlenecks and proactively improve platform efficiency and cost optimization.
- Optimize Data Warehouse and Lakehouse workloads to support high-performance reporting and analytical processing.
DevOps & Automation
- Design and implement CI/CD pipelines for Databricks solutions.
- Integrate Databricks development lifecycle with GitHub, GitHub Actions, and enterprise DevOps processes.
- Automate deployment, testing, code validation, and release management processes.
- Establish infrastructure and data engineering best practices.
Stakeholder Management
- Engage with business users, data consumers, architects, analysts, and technology leadership to gather requirements and deliver data solutions.
- Translate business requirements into scalable technical designs, data models, and platform capabilities.
- Communicate effectively with stakeholders across multiple organizational levels.
- Work independently while managing priorities and ensuring timely delivery of commitments.
- Provide technical guidance and mentorship to junior team members where required.
Production Support
- Participate in a rotating production support roster.
- Troubleshoot production incidents and prioritize issue resolution within established SLA requirements.
- Conduct root cause analysis and implement preventive measures.
- Ensure platform stability, reliability, and operational excellence.
Required Qualifications
Technical Skills
- Strong experience in:
- Bachelor's Degree in Computer Science, Information Technology, Engineering, Data Science, or a related field.
- 5-8 years of experience in Data Engineering, Data Warehousing, or Big Data technologies.
- Minimum 4+ years of hands-on Databricks experience in enterprise environments.
- Experience designing and implementing enterprise Data Warehouse solutions and modern Lakehouse.
- PySpark development and optimization.
- Delta Lake and Delta Tables.
- Unity Catalog.
- Databricks Workflows.
- RBAC and ABAC implementation within Unity Catalog.
- GitHub and Git-based development workflows.
- CI/CD implementation using GitHub Actions or equivalent.
- SQL and advanced query optimization.
- Cloud platforms (Azure, AWS, or GCP).
- Data security, governance, and compliance frameworks.
- Enterprise Data Warehouse architecture and implementation.
- Dimensional data modeling (Star Schema and Snowflake Schema).
- Fact and Dimension modeling.
- Slowly Changing Dimensions (SCD Type 1 & Type 2).
- Data Warehouse performance tuning and optimization.
Additional Technical Knowledge
- Data warehouse concepts and methodologies.
- Lakehouse architecture and Medallion design patterns.
- Data observability and monitoring.
- Infrastructure-as-Code (Terraform preferred).
Soft Skills
- Strong ownership mindset with high accountability and commitment to delivery.
- Ability to work independently with minimal supervision.
- Excellent analytical and problem-solving skills.
- Strong communication and stakeholder management capabilities.
- Ability to work effectively with stakeholders across business and technical functions.
- Ability to manage multiple priorities in a fast-paced environment.
- Collaborative team player with a proactive and customer-focused attitude.
- Willingness to participate in production support and on-call rotation schedules.
Preferred Qualifications
- Databricks Certified Data Engineer Associate or Professional certification.
- Experience implementing enterprise data governance frameworks.
Success Factors
- Experience supporting large-scale Lakehouse and Data Warehouse architectures.
- Experience working with regulated industries and compliance requirements.
- Knowledge of Data Mesh and modern data platform architectures.
- Experience integrating Databricks with Power BI, Tableau, or other BI and analytics platforms.
- The successful candidate will:
- Be the go-to Databricks engineering expert within the team.
- Drive governance and security best practices through Unity Catalog.
- Deliver reliable, scalable, and high-performing data solutions and enterprise data warehouse platforms.
- Partner effectively with business and technical stakeholders.
- Demonstrate strong ownership, accountability, and operational excellence.
- Contribute to continuous improvement of the organization's modern data platform capabilities.