freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer Consultant

Summary

Build and maintain scalable GCP-based data pipelines using BigQuery, Dataflow, and Composer; design ETL mappings and orchestrate workflows with CI/CD and monitoring.

Enterprise Data Ingestion and Data Engineering

  • Design and build scalable reusable ingestion pipelines realtime and batch on GCP GCS BigQuery Dataflow; develop parameterized pipelines using Google‑native services and or Informatica IDMC mappings and taskflows.
  • Implement CDC patterns, idempotent loads, late‑arriving data handling and schema evolution; optimize BigQuery ingestion strategies between batch and streaming, partitioning, and clustering.
  • Establish version control, CI/CD, and environment promotion for dev, test, and prod.

Results

  • GCP‑native pipelines (Dataflow, Composer, Cloud Run) and IDMC mappings/taskflows with BigQuery schema designs across raw, stage, and curated zones; implement partitioning, clustering, CI/CD pipelines, deployment artifacts, and operational runbooks.

Data Mapping & Transformation Design

  • Profile source data to analyze data quality and patterns.
  • Design field‑level mappings, business rules, joins, aggregations, and derivations.
  • Implement transformations using BigQuery SQL, Dataflow, or IDMC transformation logic.

Results

  • Source‑to‑Target Mapping documentation and transformation mappings (STM).

Orchestration, Automation & Reliability Engineering

  • Build and parameterize end‑to‑end workflows using Cloud Composer, Airflow, or IDMC taskflows.
  • Define job dependencies, schedules, SLAs, and failure‑handling strategies.
  • Implement retries, backoff strategies, checkpoints, and restartability.
  • Integrate monitoring, logging, and alerting with Cloud Monitoring, Cloud Logging, and ChatOps tools.

Results

  • Production‑ready orchestration workflows with environment‑aware parameters; scheduling calendars, dependency diagrams, and SLA matrices.
  • Alerting, monitoring dashboards, and operational SOPs.

Data Pipelines Monitoring, Performance & Cost Optimization

  • Monitor pipeline health, data freshness, volumes, and anomaly patterns; support and monitor data pipelines during off‑hours and weekends.
  • Track and manage SLAs related to runtime failures, cost, and data latency.
  • Optimize BigQuery performance through query refactoring, partition pruning, materialized views.
  • Manage and optimize GCP costs: storage lifecycle, slot usage, query optimization, caching.

Results

  • Daily and weekly operational and SLA reports, performance and cost optimization plans with before‑and‑after metrics, BigQuery tuning guidelines, and data materialization strategy.

Requirements

  • 5 years of experience as a Data Engineer or ETL Developer.
  • Strong experience with Google Cloud Platform, BigQuery, GCS, DataStream, Dataflow, Cloud Composer, IAM, Cloud Monitoring.
  • Proficiency in SQL, BigQuery‑optimized and data modeling for analytics.
  • Hands‑on experience building production‑grade scalable data pipelines.
  • Experience with CI/CD, Git‑based version control, and DevOps practices.
  • Optional: Experience with Informatica Intelligent Data Management Cloud, IDMC, CDI, Mass Ingestion, Taskflows.
  • Bachelor's degree in Computer Science, Information Technology, Information Systems, or a related field.

See also