Senior Data Engineer - Data Engineering
Summary
Design and own data pipelines and golden datasets for Plaid’s financial data platform using SQL, Python, and tools like Airflow, Snowflake, and Spark.
You will design golden datasets and data usage principles, lead data engineering projects, own SQL and Python pipelines for the data lake and warehouse, improve data quality and performance, adopt industry tools, and define dataset quality, uptime, and usefulness.
Responsibilities
- Understand product and strategy to inform dataset choices and data usage principles
- Design datasets with data quality and performance in mind
- Lead data engineering projects
- Advocate for industry tools and practices
- Own SQL and Python data pipelines
- Document data and define dataset quality, uptime, and usefulness
Requirements
- 4+ years of dedicated data engineering experience
- Experience building data models and pipelines on large datasets
- SQL
- DBT
- Mode
- Airflow
- Redshift
- Snowflake
- Databricks
- Spark
- Kafka
- Schema design
- Data privacy
- Data integrity
Benefits
- Equity