Data Engineer
Summary
Builds and maintains ETL/ELT pipelines and scalable data warehouses, lakes, and lakehouses, with a focus on data quality and query/performance tuning. Core stack spans SQL, Python/Scala/Java, Spark, Airflow, and Terraform across AWS, Azure, and GCP.
Job Description
- Pipeline Development: Build and maintain robust ETL/ELT pipelines
- Data Architecture: Design scalable data warehouses, lakes, and lakehouses
- Data Quality: Implement automated validation, cleansing, and monitoring systems
- Performance Tuning: Optimize slow-running queries and data processing jobs
- Collaboration: Work with technical teams to maintain clean datasets
PRIMARY SKILLS
- Terraform or CloudFormation, CI/CD Pipelines, or Dagster DevOps / DataOps
- Proficiency with Git, Prefect, and Data Vault Orchestration
- Experience Managing Workflows using Apache Airflow, star/snowflake schemas, or Hadoop Data Modeling
- Expertise in Kimball Methodology, Flink, Dataproc) is a plus.
- Big Data Frameworks: Hands on Experience with Apache Spark, Data Factory) / GCP (BigQuery, Redshift).
- Knowledge of Azure (Synapse, EMR, Cloud Platforms)
- Extensive Experience with AWS (Glue, and Scala or Java, SQL, Python