Data Engineer - AWS/Pyspark
Key Responsibilities
Design, develop, and maintain scalable data pipelines using
PySpark on AWS. Build and orchestrate ETL/ELT workflows using
AWS Glue and
AWS Step Functions . Develop serverless applications and automation using
AWSLambda . Write clean, efficient, and maintainable Python/PySpark codefollowing engineering best practices. Provision and manage cloud infrastructure using
Terraform (Infrastructure as Code). Implement and maintain
CI/CD pipelines
to automate codedeployment, testing, and infrastructure changes. Monitor, troubleshoot, and optimize data pipelines forperformance, reliability, and cost efficiency. Collaborate with business stakeholders to deliver datasolutions. Follow DevOps, security, and coding standards throughout theengagement.
Required Skills Strong hands-on experience with
PySpark
and Python fordata engineering. Experience developing ETL pipelines using
AWS Glue . Proficiency with
AWS Step Functions
for workfloworchestration. Experience building serverless solutions using
AWS Lambda . Hands-on experience with
Terraform
for Infrastructure asCode (IaC). Experience implementing
CI/CD
pipelines using tools suchas GitLab, GitHub Actions, Jenkins, or similar. Good understanding of AWS services, data lakes, IAM, S3,CloudWatch, and monitoring.
PySpark on AWS. Build and orchestrate ETL/ELT workflows using
AWS Glue and
AWS Step Functions . Develop serverless applications and automation using
AWSLambda . Write clean, efficient, and maintainable Python/PySpark codefollowing engineering best practices. Provision and manage cloud infrastructure using
Terraform (Infrastructure as Code). Implement and maintain
CI/CD pipelines
to automate codedeployment, testing, and infrastructure changes. Monitor, troubleshoot, and optimize data pipelines forperformance, reliability, and cost efficiency. Collaborate with business stakeholders to deliver datasolutions. Follow DevOps, security, and coding standards throughout theengagement.
Required Skills Strong hands-on experience with
PySpark
and Python fordata engineering. Experience developing ETL pipelines using
AWS Glue . Proficiency with
AWS Step Functions
for workfloworchestration. Experience building serverless solutions using
AWS Lambda . Hands-on experience with
Terraform
for Infrastructure asCode (IaC). Experience implementing
CI/CD
pipelines using tools suchas GitLab, GitHub Actions, Jenkins, or similar. Good understanding of AWS services, data lakes, IAM, S3,CloudWatch, and monitoring.