6 + YoE - Data Engineer – Big Data / PySpark - any UST Location - Immediate Joiner

Summary

Data Engineer building and optimizing Big Data pipelines with PySpark, Apache Spark, Airflow, and Python, including Spark performance tuning and CI/CD integration on cloud platforms (GCP preferred).

Candidates ready to join immediately can share their details via email for quick processing.

CCTC | ECTC | Notice Period | Location Preference

nitin.patil@ust.com

Act fast for immediate attention! ⏳

Must-Have Skills

  • 6+ years of overall experience in Data Engineering / Big Data.
  • Strong understanding of Big Data concepts and architecture.
  • Strong hands-on experience with Apache Spark.
  • Expertise in:
  • Spark Performance Tuning
  • Spark Optimization
  • Query/Job Performance Improvement
  • Troubleshooting Spark workloads
  • Strong hands-on experience with PySpark and Spark.
  • Strong programming experience in Python.
  • Good experience working with MySQL / SQL.
  • Strong experience in designing and developing Data Pipelines.
  • Hands-on experience with Apache Airflow for data pipeline orchestration and scheduling.
  • Experience working with at least one Cloud Platform.
  • GCP experience is preferred.
  • Good understanding of CI/CD and DevOps concepts.
  • Experience integrating data engineering workloads with CI/CD pipelines.
  • Strong debugging, troubleshooting, and problem-solving skills.

Preferred Skills

  • Hands-on exposure to relevant GCP data services.
  • Experience handling large-scale and high-volume datasets.
  • Understanding of distributed data processing and data architecture.
  • Experience improving the scalability, reliability, and performance of data pipelines.
  • Exposure to Agile development and DevOps practices.

Key Responsibilities

  • Design, develop, and maintain scalable Big Data and Data Engineering solutions.
  • Develop data processing applications using Python, PySpark, and Apache Spark.
  • Perform Spark performance tuning and optimization for large-scale workloads.
  • Build, maintain, and monitor robust ETL/ELT data pipelines.
  • Develop and manage workflow orchestration using Apache Airflow.
  • Work with MySQL/SQL for data extraction, transformation, and validation.
  • Deploy and support data engineering solutions in cloud environments, preferably GCP.
  • Work with DevOps teams to implement and maintain CI/CD pipelines.
  • Troubleshoot production issues and optimize data processing performance.
  • Collaborate with engineering and business teams to deliver reliable and scalable data solutions.

Primary Skill Combination

Big Data + Apache Spark + PySpark + Python + Airflow + SQL/MySQL + Cloud (GCP Preferred) + CI/CD

Mandatory Focus: Strong hands-on Apache Spark performance tuning and optimization experience.


See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available