Senior Data Engineer
Summary
Build and maintain batch and real-time data pipelines on GCP, ensuring data quality, governance, and reliability for analytics and ML workloads.
You will design, build, and operate batch and real-time data pipelines on GCP. You will implement data-quality controls and observability, maintain ingestion connectors and integrations, contribute to engineering standards, support data-science pipelines, and help ensure governance, lineage, reliability, and change control.
Responsibilities
- Design and build batch and real-time data pipelines on GCP
- Implement Kafka event streams, BigQuery transformations, and Airflow workflows
- Define and implement data contracts, quality checks, and observability
- Maintain in-house connectors and third-party ingestion integrations
- Contribute to code review, dbt modelling, CI/CD, and change-control standards
- Support feature-engineering and model-data pipelines
- Implement lineage, reproducibility, and freshness safeguards for ML workloads
Requirements
- Production data-pipeline experience on GCP
- BigQuery
- Cloud Composer or Airflow
- Google Cloud Storage
- Cloud Run
- SQL
- Python
- Apache Kafka or equivalent streaming technology
- Event-driven ingestion
- dbt
- Medallion architecture
- Data contract
- Lineage tracking
- Observability
- Pipeline SLO monitoring
- Financial-services data governance
- Access control
- Audit trail
- Change control
Benefits
- Extra time off for volunteering and community work