Data Engineer - AWS/Pyspark
Summary
Build and maintain scalable data pipelines using PySpark and AWS services like Glue, Lambda, and Step Functions.
Key Responsibilities
- Design, develop, and maintain scalable data pipelines using PySparkon AWS.
- Build and orchestrate ETL/ELT workflows using AWS Glueand AWS Step Functions.
- Develop serverless applications and automation using AWSLambda.
- Write clean, efficient, and maintainable Python/PySpark codefollowing engineering best practices.
- Provision and manage cloud infrastructure using Terraform(Infrastructure as Code).
- Implement and maintain CI/CD pipelines to automate codedeployment, testing, and infrastructure changes.
- Monitor, troubleshoot, and optimize data pipelines forperformance, reliability, and cost efficiency.
- Collaborate with business stakeholders to deliver datasolutions.
- Follow DevOps, security, and coding standards throughout theengagement.
Required Skills
- Strong hands-on experience with PySpark and Python fordata engineering.
- Experience developing ETL pipelines using AWS Glue.
- Proficiency with AWS Step Functions for workfloworchestration.
- Experience building serverless solutions using AWS Lambda.
- Hands-on experience with Terraform for Infrastructure asCode (IaC).
- Experience implementing CI/CD pipelines using tools suchas GitLab, GitHub Actions, Jenkins, or similar.
- Good understanding of AWS services, data lakes, IAM, S3,CloudWatch, and monitoring.