Pyspark Data Engineer

Summary

Design, develop, and maintain scalable ETL/ELT data pipelines using PySpark, process large datasets, build data workflows, and maintain data models, data lakes, and data warehouse solutions.

  • Design, develop, and maintain scalable ETL/ELT data pipelines using PySpark.
  • Process and transform large datasets from various structured and unstructured sources.
  • Build robust data workflows and optimize Spark jobs for performance and scalability.
  • Develop and maintain data models, data lakes, and data warehouse solutions.
  • Collaborate with Data Scientists, Analysts, and Business Teams to understand data requirements.
  • Ensure data quality, integrity, security, and governance standards are maintained.
  • Troubleshoot and resolve performance bottlenecks in Spark applications.
  • Monitor data pipelines and automate operational processes.
  • Implement CI/CD practices for data engineering workflows.
  • Create and maintain technical documentation for data pipelines and processes.

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available