Data Engineer – Microsoft Fabric
Summary
The Data Engineer will design and maintain scalable ETL/ELT pipelines and data transformation workflows using Microsoft Fabric and PySpark. The role involves working with large-scale datasets to support enterprise analytics and business intelligence initiatives.
About Applix
At Applix, we build intelligent engineering and enterprise technology solutions that help global organizations accelerate digital transformation. Our teams develop scalable cloud platforms, AI-powered applications, enterprise integrations, and data-driven solutions that enable innovation across manufacturing, automotive, healthcare, finance, and other industries.
We are looking for a Data Engineer – Microsoft Fabric with strong expertise in PySpark and Microsoft Fabric to design, develop, and optimize scalable data processing solutions that support enterprise analytics and business intelligence initiatives.
Job Summary
The Data Engineer will be responsible for designing and implementing scalable data pipelines and transformation workflows using Microsoft Fabric and PySpark. The ideal candidate will have experience building high-performance ETL/ELT solutions, working with large-scale datasets, and leveraging Fabric services such as Lakehouse, Notebooks, Spark, and Data Factory Pipelines to deliver reliable and efficient data platforms.
Key Responsibilities
- Design, develop, and maintain scalable data processing solutions using PySpark within Microsoft Fabric.
- Build and optimize ETL/ELT pipelines for processing large-scale structured and semi-structured datasets.
- Develop data ingestion, transformation, cleansing, and aggregation workflows using Fabric Notebooks, Lakehouse, and Data Factory Pipelines.
- Leverage Spark DataFrames and Spark SQL to implement efficient and high-performance data transformations.
- Design and manage data models and storage structures within Microsoft Fabric Lakehouse.
- Monitor, troubleshoot, and optimize data pipelines for performance, scalability, and reliability.
- Collaborate with data architects, BI developers, and business stakeholders to deliver analytics-ready datasets.
- Implement data quality, validation, and governance best practices throughout the data lifecycle.
- Participate in code reviews and contribute to development standards, documentation, and deployment processes.
Required Skills & Qualifications
- 5+ years of experience in Data Engineering or related roles.
- Strong hands-on experience with Microsoft Fabric and PySpark.
- Expertise in building scalable ETL/ELT pipelines using Spark DataFrames and Spark SQL.
- Experience with Microsoft Fabric Lakehouse, Fabric Notebooks, and Data Factory Pipelines.
- Strong understanding of distributed data processing concepts and performance optimization techniques.
- Experience working with structured and semi-structured data formats such as Parquet, Delta, JSON, and CSV.
- Proficiency in SQL and data transformation techniques.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent communication and collaboration abilities.
Preferred Qualifications
- Experience with Delta Lake and modern data lake architectures.
- Exposure to Azure Data Services and cloud-based analytics platforms.
- Familiarity with CI/CD practices for data engineering workflows.
- Knowledge of data governance, security, and monitoring best practices.
- Experience supporting enterprise reporting and analytics solutions.
Experience: 5+ years in data engineering with hands-on experience in PySpark, Microsoft Fabric, and modern cloud-based data platforms.