freehire launches on Product Hunt on 26 August.

Follow →

Senior Data Engineer (PySpark & Microsoft Fabric)

Summary

Build and maintain scalable data pipelines using PySpark and Microsoft Fabric to move and transform data into analytics-ready datasets.

Key Responsibilities



  • Design, build, and maintain scalable data pipelines that extract, transform, and load data from multiple sources into centralized platforms such as data warehouses, lakehouses, and operational databases.

  • Develop and manage end-to-end data workflows using Microsoft Fabric, including integration with Azure Data Factory and related Azure data services.

  • Implement physical and logical data models that support efficient storage, performance, and analytical use cases, while maintaining data integrity and reliability.

  • Transform raw and semi-structured data into analytics-ready datasets through cleansing, standardization, and deduplication processes.

  • Integrate data from disparate systems and ensure consistency, accuracy, and quality across the data pipeline.

  • Optimize database and pipeline performance through tuning, monitoring, and proactive issue resolution.

  • Collaborate with analytics, reporting, and business teams to translate data requirements into robust technical solutions.



Required Competencies



  • Bachelor’s degree in Computer Science, Information Technology, Mathematics, or a related field.

  • At least 6 years of progressive, hands-on experience in Data Engineering.

  • Minimum 6 years of experience in ETL/ELT development, data warehousing, and data modeling.

  • At least 3 years of experience delivering data solutions on the Microsoft Azure platform, with strong exposure to Microsoft Fabric and Azure Data Factory.

  • Strong proficiency in PySpark and/or Python for large-scale data processing.

  • Solid experience with T-SQL and relational database concepts.

  • Proven expertise in designing, building, and maintaining robust ETL pipelines for structured and semi-structured data in warehouse and lakehouse environments.

  • Strong understanding of data architecture, performance optimization, and best practices for scalable data platforms.

See also