Senior Data Engineer
Summary
Build and maintain scalable data pipelines and ETL workflows using Databricks, Apache Spark, and cloud services to power analytics and ML workloads.
Responsibilities
- Design, build, and maintain data pipelines and ETL processes using Databricks and Apache Spark.
- Optimize data workflows for performance, scalability, and cost efficiency.
- Implement data Lakehouse architecture and manage data ingestion from multiple sources.
- Collaborate with data scientists and analysts to enable advanced analytics and machine learning workloads.
- Ensure data quality, governance, and security across all data assets.
- Monitor and troubleshoot Databricks clusters, jobs, and workflows.
- Integrate Databricks with cloud services (AWS, Azure, or GCP) and other enterprise systems.
- Document processes, standards, and best practices for data engineering.
Requirements
- 3+ years of experience in data engineering or big data technologies.
- Hands‑on experience with Databricks, Apache Spark, and PySpark.
- Strong knowledge of SQL, Python, and data modeling principles.
- Experience with cloud platforms (AWS, Azure, or GCP) and their data services.
- Familiarity with Delta Lake, Lakehouse architecture, and data governance.
- Understanding of CI/CD pipelines and DevOps practices for data workflows.
- Excellent problem‑solving and communication skills.
Core Competencies
Demonstrates expertise in designing and maintaining data pipelines and ETL processes using Databricks and Apache Spark, with a strong focus on data quality, governance, and integration with cloud services. Proficient in optimizing workflows for performance and scalability while collaborating effectively with data scientists and analysts.