PySpark Data Engineer
Summary
Builds and optimizes data pipelines using PySpark, SQL, and the Hadoop ecosystem to process large datasets and support business analytics.
Roles and
Responsibilities:
- Responsible
for developing and maintaining applications with PySpark
- Contribute
to the overall design and architecture of the application developed and
deployed.
- Performance
Tuning wrt to executor sizing and other environmental parameters, code
optimization, partitions tuning, etc
- Interact
with business users to understand requirements and troubleshoot issues.
- Implement
Projects based on functional specifications.
Must-Have
Skills:
- Relevant
Experience: 3-6 Years
- SQL -
Mandatory
- Python -
Mandatory
- SparkSQL -
Mandatory
- PySpark -
Mandatory
- Hive -
Mandatory
- HDFS and
Spark - Mandatory
- Scala -
Advantage
- Apache
Airflow - Advantage