Cloudera Data Engineer
Role: Data Migration Specialist.
Experience: 6+ years' experience in large dataplatform and large scale data migration in Cloudera ADF & Databricks
Key Responsibilities (R&R) • Own end to end historical data migration from Cloudera (Hive/HDFS) and legacy systems • Design and execute one off bulk history migration • Design daily incremental loads via ADF, fully metadata driven • Implement migration using customer ingestion framework or custom framework • Design restartable, idempotent migration pipelines • Build reconciliation reports (counts, checksums, balances) • Develop automated testing packs for HDM (history data migration) • Handle schema drift, data quality issues, and legacy anomalies • Support parallel run, cutover, and business sign off • Act as migration SPOC across platform, ingestion, and business teams
Mandatory Skills (Must Have) • Azure Data Factory (ADF) – metadata driven pipelines • Basic Cloudera working knowledge • Cloudera
Cloud migration (mandatory) • Large volume historical data migration • Azure Databricks (Spark / SQL) • Data reconciliation and audit sign off • Framework based data ingestion / migration
Good to Have • Azure Data Box–based migration experience • Bulk history extraction using Data Box • Landing data into ADLS Gen2 • Processing history in Databricks (Delta) • Orchestrating incrementals via ADF
Experience: 6+ years' experience in large dataplatform and large scale data migration in Cloudera ADF & Databricks
Key Responsibilities (R&R) • Own end to end historical data migration from Cloudera (Hive/HDFS) and legacy systems • Design and execute one off bulk history migration • Design daily incremental loads via ADF, fully metadata driven • Implement migration using customer ingestion framework or custom framework • Design restartable, idempotent migration pipelines • Build reconciliation reports (counts, checksums, balances) • Develop automated testing packs for HDM (history data migration) • Handle schema drift, data quality issues, and legacy anomalies • Support parallel run, cutover, and business sign off • Act as migration SPOC across platform, ingestion, and business teams
Mandatory Skills (Must Have) • Azure Data Factory (ADF) – metadata driven pipelines • Basic Cloudera working knowledge • Cloudera
Cloud migration (mandatory) • Large volume historical data migration • Azure Databricks (Spark / SQL) • Data reconciliation and audit sign off • Framework based data ingestion / migration
Good to Have • Azure Data Box–based migration experience • Bulk history extraction using Data Box • Landing data into ADLS Gen2 • Processing history in Databricks (Delta) • Orchestrating incrementals via ADF