Data Engineer Consultant
Summary
Build and maintain scalable GCP-based data pipelines using BigQuery, Dataflow, and Composer; design ETL mappings and orchestrate workflows with CI/CD and monitoring.
Enterprise Data Ingestion and Data Engineering
- Design and build scalable reusable ingestion pipelines realtime and batch on GCP GCS BigQuery Dataflow; develop parameterized pipelines using Google‑native services and or Informatica IDMC mappings and taskflows.
- Implement CDC patterns, idempotent loads, late‑arriving data handling and schema evolution; optimize BigQuery ingestion strategies between batch and streaming, partitioning, and clustering.
- Establish version control, CI/CD, and environment promotion for dev, test, and prod.
Results
- GCP‑native pipelines (Dataflow, Composer, Cloud Run) and IDMC mappings/taskflows with BigQuery schema designs across raw, stage, and curated zones; implement partitioning, clustering, CI/CD pipelines, deployment artifacts, and operational runbooks.
Data Mapping & Transformation Design
- Profile source data to analyze data quality and patterns.
- Design field‑level mappings, business rules, joins, aggregations, and derivations.
- Implement transformations using BigQuery SQL, Dataflow, or IDMC transformation logic.
Results
- Source‑to‑Target Mapping documentation and transformation mappings (STM).
Orchestration, Automation & Reliability Engineering
- Build and parameterize end‑to‑end workflows using Cloud Composer, Airflow, or IDMC taskflows.
- Define job dependencies, schedules, SLAs, and failure‑handling strategies.
- Implement retries, backoff strategies, checkpoints, and restartability.
- Integrate monitoring, logging, and alerting with Cloud Monitoring, Cloud Logging, and ChatOps tools.
Results
- Production‑ready orchestration workflows with environment‑aware parameters; scheduling calendars, dependency diagrams, and SLA matrices.
- Alerting, monitoring dashboards, and operational SOPs.
Data Pipelines Monitoring, Performance & Cost Optimization
- Monitor pipeline health, data freshness, volumes, and anomaly patterns; support and monitor data pipelines during off‑hours and weekends.
- Track and manage SLAs related to runtime failures, cost, and data latency.
- Optimize BigQuery performance through query refactoring, partition pruning, materialized views.
- Manage and optimize GCP costs: storage lifecycle, slot usage, query optimization, caching.
Results
- Daily and weekly operational and SLA reports, performance and cost optimization plans with before‑and‑after metrics, BigQuery tuning guidelines, and data materialization strategy.
Requirements
- 5 years of experience as a Data Engineer or ETL Developer.
- Strong experience with Google Cloud Platform, BigQuery, GCS, DataStream, Dataflow, Cloud Composer, IAM, Cloud Monitoring.
- Proficiency in SQL, BigQuery‑optimized and data modeling for analytics.
- Hands‑on experience building production‑grade scalable data pipelines.
- Experience with CI/CD, Git‑based version control, and DevOps practices.
- Optional: Experience with Informatica Intelligent Data Management Cloud, IDMC, CDI, Mass Ingestion, Taskflows.
- Bachelor's degree in Computer Science, Information Technology, Information Systems, or a related field.