Senior Data Engineer
Summary
Data Engineer building scalable, high-performance data solutions on Google Cloud Platform, including ETL/ELT pipelines, BigQuery optimization, and workflow orchestration using Cloud Composer and Dataflow.
As a Data Engineer, you’ll play a key role in building modern, scalable, and high-performance data solutions on Google Cloud Platform (GCP). You’ll be part of our growing Data & AI team, designing and implementing data architectures that help clients unlock the full potential of their data.
Your job’s key responsibilities are:
Building efficient and scalable ETL/ELT processes to ingest, transform, and load data from various structured and unstructured sources (databases, APIs, streaming platforms) into BigQuery and Cloud Storage
Implementing data ingestion and real-time processing using Dataflow (Apache Beam) and Pub/Sub for batch and streaming workflows
Developing SQL transformation workflows with Dataform, including version control, testing, and automated scheduling with built-in quality assertions
Creating efficient, cost-optimized BigQuery queries with proper partitioning, clustering, and denormalization strategies
Orchestrating complex workflows using Cloud Composer (Apache Airflow) and Cloud Functions for event-driven data processing
Implementing centralized data governance and metadata management using Dataplex with automated cataloging and lineage tracking
Monitoring and optimizing data pipelines for performance, scalability, and cost using Cloud Monitoring and Cloud Logging
Collaborating with data scientists and analysts to understand data requirements and deliver actionable insights
Staying up to date with GCP advancements in data services, BigQuery features, and data engineering best practices
Essential Skills:
3+ years of hands-on experience as a Data Engineer with proven expertise in Google Cloud Platform (GCP)
Strong experience with BigQuery (SQL, partitioning, clustering, optimization) and Dataflow (Apache Beam)
Strong programming skills in Python with experience in data manipulation libraries (PySpark, pandas)
Expert-level SQL proficiency for complex transformations, optimization, and analysis
Proficiency with Dataform for modular SQL-based data transformations and data pipeline management
Solid understanding of data warehousing principles, ETL/ELT processes, dimensional modeling, and data governance
Experience integrating data from various APIs and streaming systems (Pub/Sub)
Cloud Composer experience for workflow orchestration
Excellent communication and collaboration skills in English (min. B2 level)
Ability to work independently and as part of an agile team
Beneficial Skills:
Google Professional Data Engineer certification
Knowledge of BigLake for unified access and management of structured and unstructured data
Experience with Dataplex for managing metadata, lineage, and data governance
Familiarity with Infrastructure-as-Code (Terraform) for automating GCP resource provisioning and CI/CD pipelines
Experience with data visualization tools such as Looker, Looker Studio, or Power BI
Interest in or experience with machine learning workflows using Vertex AI or similar platforms