Staff Software Engineer - Data Infrastructure
Summary
Build and maintain Plaid’s data infrastructure, including warehouses, lakehouses, Spark, and streaming systems, to support ML and analytics workflows.
You will shape data infrastructure roadmaps and deliver systems for data warehouses, lakehouses, Spark, workflow orchestration, and streaming. You will improve machine-learning development paths and data freshness, reduce operational burden, collaborate across functions, and mentor engineers through technical reviews and guidance.
Responsibilities
- Contribute to the long-term roadmap for data-driven and machine-learning iteration
- Lead data infrastructure projects involving ML development paths, streaming, ETL, warehouses, and lakehouses
- Define technical roadmaps for backend systems and abstractions with stakeholders
- Debug and troubleshoot the Data Platform
- Reduce operational burden
- Mentor engineers and review technical documents and code changes
Requirements
- 6+ years of software engineering experience
- Hands-on software engineering experience delivering projects in data infrastructure or platform domains
- Deep understanding of data warehouses, data lakehouses, Apache Spark, streaming infrastructure, or workflow orchestration
- Strong cross-functional collaboration and communication skills
- Project management skills
- Proficiency in coding, testing, and system design
- Experience mentoring and guiding junior engineers
- Experience with Databricks, Airflow, AWS EMR, or Python
Benefits
- Equity