Data Migration Engineer
We are seeking a detail-oriented and capable Data Migration Engineer to join our Data & AI practice. The successful candidate will bring solid experience in data migration, ETL/ELT pipeline development, and cloud-based data platforms, with a focus on AWS Data Lakehouse environments.
This role is key to supporting the design, build, and validation of data migration pipelines, enabling the successful transition of data from legacy systems to modern cloud platforms. You will contribute to ensuring data quality, integrity, and performance, particularly through structured testing and validation activities.
You will work closely with architects, senior engineers, and analysts to deliver scalable and reliable migration solutions, using technologies such as AWS Glue, Apache Iceberg, Python/PySpark, SQL, and YAML configurations. You should be comfortable working in a collaborative, delivery-focused environment and have a strong interest in data migration, cloud technologies, and modern data engineering practices.
What youll be doing:
Client Engagement & Delivery
• Support delivery within data migration programmes, contributing to key workstreams
• Collaborate with architects, engineers, and stakeholders to implement migration solutions
• Assist in planning and executing data migration tasks and deliverables
Data Migration Engineering
• Build and maintain data migration pipelines from legacy data warehouses to AWS-based platforms
• Develop ETL/ELT pipelines using:
• AWS Glue
• Python / PySpark
• SQL
• YAML configurations
• Support execution of bulk data migrations and incremental/delta loads
• Assist with pipeline repointing and migration to cloud environments
Data Pipeline Testing & Validation (Core Focus)
• Test ETL/ELT data pipelines on AWS services, including AWS Glue and Apache Iceberg
• Support validation of data pipeline migrations to AWS Data Lakehouse architectures
• Test pipelines using:
• Python/PySpark transformations
• SQL-based validation logic
• YAML-driven configurations
• Execute and validate:
• Initial bulk data loads
• Incremental/delta data processing
• Write and run SQL queries to validate:
• Data completeness
• Data accuracy
• Transformation outputs
• Support development of test scripts and validation checks
AWS Data Platforms & Lakehouse
• Work with AWS services including:
• AWS Glue
• S3-based data lakes
• Support implementation of Data Lakehouse architectures, including Apache Iceberg
• Contribute to improving pipeline performance and reliability
Data Transformation & Support
• Apply transformation logic based on defined data mapping rules
• Support preparation of data for target-state models
• Assist in ensuring consistency between source and target datasets
Collaboration & Best Practices
• Work collaboratively with:
• Solution Architects
• Data Engineers
• Data Migration Architects
• Analysts and QA teams
• Follow established engineering standards and best practices
• Contribute to documentation and reusable components
Quality, Governance & Security
• Support maintenance of data quality and integrity during migration
• Follow secure data handling practices
• Assist with compliance requirements, including:
• GDPR
• Public sector data standards (where applicable)
• Contribute to testing, validation, and audit activities
What experience youll bring:
• Experience in data engineering or data migration delivery
• Strong focus on testing, validation, and data quality assurance
• Ability to work across data pipelines and transformation workflows
• Good analytical and problem-solving skills
• Effective communication and teamwork skills
• Willingness to learn and develop in data migration and cloud technologies
Technical Expertise
• Hands-on experience with:
• AWS cloud services, especially AWS Glue
• Python / PySpark
• SQL querying and validation
• YAML configuration (desirable)
• Experience testing or supporting:
• ETL/ELT pipelines
• Data migration processes
• Familiarity with:
• Data lake / Lakehouse concepts (e.g., Apache Iceberg)
• Distributed processing frameworks (e.g., Spark)
• Basic understanding of:
• ETL vs ELT approaches
• Cloud-based data architectures
• Exposure to version control and CI/CD tools desirable