Information Technology Consultant (Data Engineer (Cloudera to AWS Migration)
Summary
Design and migrate data pipelines from Cloudera/Hadoop to AWS using Python, PySpark, and AWS services like S3, Glue, EMR, and Redshift to support analytics and reporting.
We are seeking a highly skilled Data Engineer with 5+ years of experience in building scalable data platforms and modern data pipelines. The ideal candidate will have strong expertise in Python, SQL, PySpark, Hadoop/Cloudera technologies, and AWS data services, with hands‑on experience migrating enterprise data workloads from Cloudera/Hadoop environments to AWS. This role will play a key part in designing, developing, and optimizing cloud‑native data solutions that support analytics, reporting, and business-critical applications.
- Design, develop, and maintain scalable ETL processes and data pipelines using Python, SQL, and PySpark.
- Lead and support migration of data workloads, datasets, and processing frameworks from Cloudera/Hadoop environments to AWS.
- Build and optimize cloud‑native data solutions leveraging AWS services such as S3, Glue, EMR, Athena, and Redshift.
- Develop robust data ingestion, transformation, and integration frameworks to support enterprise analytics and reporting needs.
- Monitor and improve data pipeline performance, reliability, and scalability across on‑premises and cloud environments.
- Ensure data quality, governance, security, and compliance standards are implemented throughout the data lifecycle.
- Collaborate with data architects, business stakeholders, analysts, and engineering teams to translate business requirements into technical solutions.
- Implement data orchestration, automation, and operational best practices to improve platform efficiency and supportability.
- Support production deployments, troubleshooting, monitoring, and continuous improvement initiatives.
Required Skills
- 5+ years of experience in Data Engineering, Data Warehousing, or Big Data technologies.
- Strong programming expertise in Python, SQL, and PySpark.
- Hands‑on experience with the Hadoop/Cloudera ecosystem, including:
- HDFS
- Hive
- Impala
- Spark
- Cloudera Data Platform (CDP) preferred
- Strong knowledge of AWS Data Services including:
- Amazon S3
- AWS Glue
- EMR
- Athena
- Redshift
- Proven experience migrating data platforms and workloads from Cloudera/Hadoop to AWS.
- Extensive experience in ETL development, data integration, and large‑scale data pipeline implementation.
- Strong understanding of data processing, performance tuning, and optimization techniques.
- Experience working in Agile delivery environments.
Preferred Skills
- Data Modeling and Data Architecture concepts.
- Apache Airflow for workflow orchestration.
- DevOps and CI/CD pipeline implementation.
- Banking and Financial Services domain experience.
- Experience with cloud migration, modernization, and data transformation initiatives.