Data Engineer / Data Migration Analyst - Banking
Summary
Data Engineer to migrate legacy banking data to new enterprise data marts, using SQL, Python, SAS, Hive, Impala, and Oozie while validating outputs and documenting mappings.
We are seeking for a contract resources (6 month duration) to support a strategic data migration and reporting transformation initiative. The objective of the project is to migrate existing tables, datasets, and reports currently sourced from a legacy datamart and manual data sources to newly established enterprise data marts
The successful candidates will be responsible for analysing existing data processing and reporting logic, mapping legacy and manual data sources to the new data marts, redeveloping scripts and workflows, validating migrated outputs, and producing comprehensive data documentation. This role requires strong SAS, SQL and Python development skills, data analysis capabilities, and experience working with Hadoop-based data platforms
Job responsibilities:
Data Analysis & Discovery
Analyse existing tables, reports, datasets, and data processing workflows currently sourced from legacy data marts and manual processes
Review and understand existing scripts, transformation logic, data dependencies, and business rules
Identify data sources, data lineage, reporting requirements, and current-state processes
Engage business to understand existing reporting methodologies and requirements
Source Mapping & Migration
Perform source-to-target mapping between legacy data marts/manual sources and new data marts
Analyse and document field mappings, transformation rules, calculations, and business logic
Identify gaps, inconsistencies, and opportunities for process optimisation during migration
Develop migration specifications and ensure alignment with business requirements
Development & Automation
Develop, enhance, and maintain SQL scripts using Hive and Impala engines via Hue
Develop scripts for data extraction, transformation, validation, and reporting processes
Rebuild existing tables, reporting datasets, and data preparation processes using the new data marts as source systems
Configure and support workflow scheduling and automation using Oozie
Ensure developed solutions are scalable, efficient, and maintainable
Data Validation & Testing
Perform reconciliation and validation between outputs generated from current sources and outputs generated from the new data marts
Validate data accuracy, completeness, consistency, and timeliness
Investigate and resolve data discrepancies and defects identified during testing
Support User Acceptance Testing (UAT) and business validation activities
Ensure migrated reports and datasets meet defined business and technical requirements
Documentation & Governance
Produce and maintain project artefacts, including:
Source-to-target mapping documents
Data lineage documentation
Business and technical logic documentation
Metadata documentation
Data dictionaries
Reconciliation and validation reports
Technical specifications and operational runbooks
Ensure documentation complies with enterprise data governance and documentation standards
Job requirement:
Strong SQL, SAS or Python scripting and data manipulation skills.
Hands‑on experience with: Hue, Hive, Impala, Oozie
Experience in data warehousing, ETL development, data migration, or data engineering projects.
Experience working with large and complex datasets.
Strong understanding of relational databases and data modelling concepts.
Qualifications:
Degree in Computer Science, Information Systems, Engineering, Mathematics, Statistics, or a related discipline
Minimum 3 years of relevant experience in data engineering, ETL development, data warehousing, reporting, or data migration projects
Data Management Skills:
Source-to-target mapping
Data lineage analysis and documentation
Metadata management
Data dictionary creation and maintenance
Data quality assessment and reconciliation
Technical and functional documentation
Soft Skills:
Strong analytical and problem‑solving abilities
Ability to understand and reverse‑engineer legacy code and reporting processes
Good stakeholder engagement and communication skills
Ability to work independently and manage multiple deliverables within tight timelines
Preferred Experience:
Experience in banking, financial services, or large enterprise environments
Experience working with Hadoop ecosystem technologies
Exposure to reporting transformation, data modernisation, or data migration programmes
Familiarity with enterprise data governance and metadata management practices