freehire launches on Product Hunt on 26 August.

Follow →

GCP Data Engineer

Open 20d posting dated 6 days ago

Summary

Build and optimize ETL/ELT pipelines using PySpark and Python on Google Cloud Platform, focusing on BigQuery, Dataflow, and Dataproc for scalable data processing.

Required Skills

  • Strong experience in Python programming.
  • Hands-on expertise with PySpark for large-scale data processing.
  • Experience with Google Cloud Platform (GCP) services:
  • BigQuery
  • Dataproc
  • Dataflow
  • Cloud Storage
  • Pub/Sub
  • Cloud Composer (Airflow)
  • Strong SQL skills and experience with relational databases.
  • Experience building batch and real-time data pipelines.
  • Good understanding of Data Lake and Data Warehouse concepts.
  • Knowledge of distributed computing and Spark optimization techniques.
  • Experience with Git, CI/CD pipelines, and Agile methodologies.

Key Responsibilities

  • Design, develop, and maintain scalable ETL/ELT pipelines using PySpark and Python.
  • Build and manage data processing solutions on Google Cloud Platform (GCP).
  • Develop data ingestion frameworks from structured, semi-structured, and unstructured data sources.
  • Implement data transformation, cleansing, and aggregation processes for analytics and reporting.
  • Work with BigQuery, Cloud Storage, Dataproc, Dataflow, Pub/Sub, Composer (Airflow), and other GCP services.
  • Optimize PySpark jobs for performance, scalability, and cost efficiency.
  • Collaborate with Data Scientists, Analysts, and Application teams to deliver high-quality data solutions.
  • Ensure data quality, governance, security, and compliance standards are met.
  • Troubleshoot and resolve production data issues and performance bottlenecks.
  • Participate in code reviews, CI/CD implementation, and DevOps practices.

referred Skills

  • Experience with Kafka and streaming architectures.
  • Knowledge of Terraform or Infrastructure as Code (IaC).
  • Experience with Docker and Kubernetes.
  • Understanding of data modeling and dimensional modeling concepts.
  • Exposure to machine learning data pipelines.