AWS Data Engineer (Pyspark)_Indore,Pune,Hyderabad,,_6 to 9 yrs

Summary

Design and optimize scalable AWS ETL pipelines using PySpark, orchestrate workflows with Glue/MWAA/Step Functions, and manage infrastructure via Terraform and Docker/EKS.

Key Responsibilities

ETL & Pipeline Development: Design and optimize scalable ETL batch pipelines in AWS for high performance and reliability.

Orchestration: Manage data workflows using tools like AWS Glue, MWAA, or Step Functions.

Data Processing: Develop large-scale processing jobs using PySpark while ensuring data quality and integrity.

Infrastructure as Code (IaC): Automate and manage AWS infrastructure using Terraform.

Containerization: Deploy and manage applications using Amazon EKS and Docker.


Required Skills & Qualifications

AWS Services: Hands-on expertise with Glue, S3, IAM, KMS, SNS, Athena, Lambda, SQS, CloudWatch, and EC2.

Programming: Proficiency in PySpark for complex data transformations.

Migration: Proven experience moving data from on-premises systems to AWS.

DevOps Tools: Skilled in Terraform for IaC and Docker/Containers for application packaging.

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available