Data Engineer
Summary
Build and scale data pipelines on Databricks and AWS, designing robust ETL/ELT workflows with Apache Spark to power analytics and decision-making.
We're looking for a skilled Data Engineer to join our growing data platform team. In this role, you'll be at the heart of building and scaling the infrastructure that powers our data-driven decisions – designing robust pipelines, shaping our lakehouse architecture, and ensuring data flows reliably from source to insight.
You will work hands‑on with Databricks and Apache Spark on AWS, taking ownership of end‑to‑end pipeline development while collaborating closely with analytics, engineering, and product teams. If you take pride in clean, performant code and enjoy solving complex data challenges at scale, we'd love to hear from you.
Key Responsibilities
- Design, build, and maintain scalable ETL/ELT pipelines using Databricks (Apache Spark).
- Implement batch and streaming data ingestion from multiple sources (databases, APIs, event streams).
- Ensure pipelines are fault‑tolerant, efficient, and cost‑optimized on AWS.
- Develop and maintain data lake / lakehouse architectures using AWS S3, Delta Lake, and Databricks.
- Implement medallion architecture (Bronze / Silver / Gold layers).
- Optimize data storage formats (Parquet, Delta) and partitioning strategies.
- Integrate AWS services: S3 (data lake storage), IAM (access control), Lambda / Step Functions (orchestration) and manage secure data access across AWS accounts and environments.
- Develop Databricks notebooks and jobs in PySpark / SQL / Python, optimize Spark jobs for performance and cost, manage Databricks Workflows, jobs, and cluster configurations, and implement Unity Catalog for governance and data access control (if used).
Requirements
- Hands‑on experience with Databricks and Apache Spark (PySpark and/or Scala).
- Strong proficiency in Spark SQL and performance optimization techniques.
- Experience with Delta Lake and lakehouse architectures.
- Proficiency in Python for data engineering and SQL for data transformation.
- Experience writing clean, testable, and maintainable code; familiarity with Git and version control workflows.
- Strong problem‑solving and analytical thinking skills; ability to work independently and take ownership of data pipelines.
- Good communication skills and ability to collaborate with cross‑functional teams.
- Attention to detail and focus on data correctness and reliability.
Nice to Have
- Familiarity with Unity Catalog or other data governance tools.
- Experience supporting BI and analytics use cases.
- Knowledge of cost optimization in AWS and Databricks.