freehire launches on Product Hunt on 26 August.

Follow →

Information Technology Consultant (Data Engineer (Cloudera to AWS Migration)

Summary

Design and migrate data pipelines from Cloudera/Hadoop to AWS using Python, PySpark, and AWS services like S3, Glue, EMR, and Redshift to support analytics and reporting.

We are seeking a highly skilled Data Engineer with 5+ years of experience in building scalable data platforms and modern data pipelines. The ideal candidate will have strong expertise in Python, SQL, PySpark, Hadoop/Cloudera technologies, and AWS data services, with hands‑on experience migrating enterprise data workloads from Cloudera/Hadoop environments to AWS. This role will play a key part in designing, developing, and optimizing cloud‑native data solutions that support analytics, reporting, and business-critical applications.

  • Design, develop, and maintain scalable ETL processes and data pipelines using Python, SQL, and PySpark.
  • Lead and support migration of data workloads, datasets, and processing frameworks from Cloudera/Hadoop environments to AWS.
  • Build and optimize cloud‑native data solutions leveraging AWS services such as S3, Glue, EMR, Athena, and Redshift.
  • Develop robust data ingestion, transformation, and integration frameworks to support enterprise analytics and reporting needs.
  • Monitor and improve data pipeline performance, reliability, and scalability across on‑premises and cloud environments.
  • Ensure data quality, governance, security, and compliance standards are implemented throughout the data lifecycle.
  • Collaborate with data architects, business stakeholders, analysts, and engineering teams to translate business requirements into technical solutions.
  • Implement data orchestration, automation, and operational best practices to improve platform efficiency and supportability.
  • Support production deployments, troubleshooting, monitoring, and continuous improvement initiatives.

Required Skills

  • 5+ years of experience in Data Engineering, Data Warehousing, or Big Data technologies.
  • Strong programming expertise in Python, SQL, and PySpark.
  • Hands‑on experience with the Hadoop/Cloudera ecosystem, including:
  • HDFS
  • Hive
  • Impala
  • Spark
  • Cloudera Data Platform (CDP) preferred
  • Strong knowledge of AWS Data Services including:
  • Amazon S3
  • AWS Glue
  • EMR
  • Athena
  • Redshift
  • Proven experience migrating data platforms and workloads from Cloudera/Hadoop to AWS.
  • Extensive experience in ETL development, data integration, and large‑scale data pipeline implementation.
  • Strong understanding of data processing, performance tuning, and optimization techniques.
  • Experience working in Agile delivery environments.

Preferred Skills

  • Data Modeling and Data Architecture concepts.
  • Apache Airflow for workflow orchestration.
  • DevOps and CI/CD pipeline implementation.
  • Banking and Financial Services domain experience.
  • Experience with cloud migration, modernization, and data transformation initiatives.

See also