Senior Data Engineer (PySpark & Microsoft Fabric)
Summary
Build and maintain scalable data pipelines using PySpark and Microsoft Fabric to move and transform data into analytics-ready datasets.
Key Responsibilities
- Design, build, and maintain scalable data pipelines that extract, transform, and load data from multiple sources into centralized platforms such as data warehouses, lakehouses, and operational databases.
- Develop and manage end-to-end data workflows using Microsoft Fabric, including integration with Azure Data Factory and related Azure data services.
- Implement physical and logical data models that support efficient storage, performance, and analytical use cases, while maintaining data integrity and reliability.
- Transform raw and semi-structured data into analytics-ready datasets through cleansing, standardization, and deduplication processes.
- Integrate data from disparate systems and ensure consistency, accuracy, and quality across the data pipeline.
- Optimize database and pipeline performance through tuning, monitoring, and proactive issue resolution.
- Collaborate with analytics, reporting, and business teams to translate data requirements into robust technical solutions.
Required Competencies
- Bachelor’s degree in Computer Science, Information Technology, Mathematics, or a related field.
- At least 6 years of progressive, hands-on experience in Data Engineering.
- Minimum 6 years of experience in ETL/ELT development, data warehousing, and data modeling.
- At least 3 years of experience delivering data solutions on the Microsoft Azure platform, with strong exposure to Microsoft Fabric and Azure Data Factory.
- Strong proficiency in PySpark and/or Python for large-scale data processing.
- Solid experience with T-SQL and relational database concepts.
- Proven expertise in designing, building, and maintaining robust ETL pipelines for structured and semi-structured data in warehouse and lakehouse environments.
- Strong understanding of data architecture, performance optimization, and best practices for scalable data platforms.