freehire launches on Product Hunt on 26 August.

Follow →

Data Engineer (ETL & PySpark)

Summary

Designs and builds scalable ETL pipelines using Python, PySpark, and SQL to process large datasets and power analytics and reporting systems.

Job Summary

We are looking for a skilled Data Engineer with strong expertise in Python, PySpark, and SQL to design, develop, and optimize scalable data pipelines and ETL processes. The ideal candidate should have experience working with large-scale datasets, distributed computing frameworks, and cloud-based data platforms while ensuring data quality, reliability, and performance.

Key Responsibilities

Design, develop, and maintain scalable ETL/ELT data pipelines using Python and PySpark. Develop high-performance data processing solutions for structured and unstructured data. Write optimized SQL queries, stored procedures, and data transformations. Build and maintain data models, data marts, and data warehouses. Perform data cleansing, validation, and quality checks. Optimize Spark jobs for performance, scalability, and resource utilization. Integrate data from multiple sources including APIs, databases, and cloud storage. Collaborate with Data Analysts, Data Scientists, and Business teams to deliver data solutions. Troubleshoot production issues and perform root cause analysis. Ensure data governance, security, and best engineering practices. Participate in code reviews and maintain technical documentation.

Required Skills

Strong experience in Python programming. Hands-on experience with PySpark and Apache Spark. Strong SQL skills with query optimization. Experience with ETL/ELT pipeline development. Good understanding of data warehousing concepts. Experience with relational databases (Oracle, SQL Server, PostgreSQL, MySQL, etc.). Knowledge of Linux/Unix environment and Shell Scripting. Familiarity with Git and CI/CD processes. Strong analytical and problem-solving skills.

Preferred Skills

Experience with cloud platforms such as AWS, Azure, or GCP. Knowledge of Databricks, AWS Glue, or EMR. Experience with Airflow or other workflow orchestration tools. Understanding of Delta Lake, Hive, or Hadoop ecosystem. Exposure to Kafka or other streaming technologies. Knowledge of Docker and Kubernetes. Experience with Agile/Scrum methodology.

Qualifications

Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.

Nice to Have

Experience with Snowflake or Databricks. Cloud certifications (AWS, Azure, or GCP). Knowledge of DevOps practices and CI/CD pipelines.

Key Competencies

Python Development PySpark Apache Spark SQL & Query Optimization ETL/ELT Development Data Warehousing Performance Tuning Data Modeling Cloud Data Engineering Problem Solving Team Collaboration Communication Skills

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available