Senior Data Engineer (AWS)
Posted
Summary
A hands-on senior data engineer who designs, builds, and maintains scalable data pipelines on AWS using PySpark, Glue, Step Functions, Lambda, Terraform, and CI/CD, ideally with banking/financial data such as ledger and regulatory datasets.
This role requires strong data Engineering& analysis expertise across
SQL, Data modelling, DBT, Airflow,
and Lakehouse/data lake architecture, with hands-on experience querying large-scale distributed datasets directly from data lakes. Banking/financial services experience, exposure to ledger/ transactional/ regulatory data, support for migration/modernization program, familiarity, with data governance / metadata. regulatory controls. Key Responsibilities
Design, develop, and maintain scalable data pipelines using PySpark on AWS. Build and orchestrate ETL/ELT workflows using AWS Glue and AWS Step Functions. Develop serverless applications and automation using AWS Lambda. Write clean, efficient, and maintainable Python/PySpark code following engineering best practices. Provision and manage cloud infrastructure using Terraform (Infrastructure as Code). Implement and maintain CI/CD pipelines to automate code deployment, testing, and infrastructure changes. Monitor, troubleshoot, and optimize data pipelines for performance, reliability, and cost efficiency. Collaborate with business stakeholders to deliver data solutions. Follow DevOps, security, and coding standards throughout the engagement. Required Skills
Strong hands-on experience with PySpark and
Python
for data engineering. Experience developing ETL pipelines using AWS Glue. Proficiency with AWS Step Functions for workflow orchestration. Experience building serverless solutions using AWS Lambda. Hands-on experience with Terraform for Infrastructure as Code (IaC). Experience implementing CI/CD pipelines using tools such as GitLab, GitHub Actions, Jenkins, or similar. Good understanding of AWS services, data lakes, IAM, S3, CloudWatch, and monitoring.
SQL, Data modelling, DBT, Airflow,
and Lakehouse/data lake architecture, with hands-on experience querying large-scale distributed datasets directly from data lakes. Banking/financial services experience, exposure to ledger/ transactional/ regulatory data, support for migration/modernization program, familiarity, with data governance / metadata. regulatory controls. Key Responsibilities
Design, develop, and maintain scalable data pipelines using PySpark on AWS. Build and orchestrate ETL/ELT workflows using AWS Glue and AWS Step Functions. Develop serverless applications and automation using AWS Lambda. Write clean, efficient, and maintainable Python/PySpark code following engineering best practices. Provision and manage cloud infrastructure using Terraform (Infrastructure as Code). Implement and maintain CI/CD pipelines to automate code deployment, testing, and infrastructure changes. Monitor, troubleshoot, and optimize data pipelines for performance, reliability, and cost efficiency. Collaborate with business stakeholders to deliver data solutions. Follow DevOps, security, and coding standards throughout the engagement. Required Skills
Strong hands-on experience with PySpark and
Python
for data engineering. Experience developing ETL pipelines using AWS Glue. Proficiency with AWS Step Functions for workflow orchestration. Experience building serverless solutions using AWS Lambda. Hands-on experience with Terraform for Infrastructure as Code (IaC). Experience implementing CI/CD pipelines using tools such as GitLab, GitHub Actions, Jenkins, or similar. Good understanding of AWS services, data lakes, IAM, S3, CloudWatch, and monitoring.