PySpark Data Engineer - ETL & Spark Pipelines
Summary
Designs and maintains large-scale data pipelines using PySpark on distributed platforms like Hadoop/Hive and cloud storage, with a focus on data quality, ETL/ELT workflow design, and Spark performance tuning, while automating workflows with tools such as Airflow or Oozie.
V2 Solutions is seeking a data analytics professional to design and maintain large-scale data pipelines using PySpark. You'll work with distributed data platforms like Hadoop/Hive and cloud storage, turning raw data into actionable insights for business teams.
The role emphasizes data quality and governance, ETL/ELT workflow design, and performance tuning of Spark jobs. You will collaborate with engineers and analysts, write clean code, and automate workflows with tools such as Airflow or Oozie.