Data Engineer
Join Our Team at Lean Solutions Group (LSG)!
Lean Solutions Group (LSG) is a next-generation solutions provider combining AI-driven automation, industry expertise, and tech-powered talent. Built in the demanding Supply Chain sector, our model now supports 600+ clients across multiple industries, powered by 10,000+ employees in five countries. We help businesses achieve immediate efficiency, long-term resilience, and scalable growth by integrating intelligent technology, optimized processes, and high-performance teams.
At LSG, we believe in your talent and your potential. Join a multicultural, people-first environment where you can grow, sharpen your skills, and unlock new career opportunities. Here, every day brings fresh challenges, collaboration, and purpose.
Our Mission:
Transform business challenges into lasting success through purpose-built teams, technology, and expertise.
Our Vision:
A world where people, empowered by technology, turn any challenge into a catalyst for growth.
About the Role
We are seeking a Data Engineer who combines solid engineering discipline, modern lakehouse expertise, and a clear understanding of how operational data becomes trusted reporting. In this role, you will build and operate the data foundation behind enterprise Business Intelligence, from ingestion pipelines and lakehouse layers to the governed, reusable datasets our analysts, semantic models, and AI tools depend on.
This position goes beyond moving data. You will replace manual uploads with automated, monitored pipelines, model data so one definition serves every report, and ensure every dataset is traceable, tested, and secure.
Key Responsibilities
Lakehouse Platform Engineering: Build and operate the enterprise lakehouse on Microsoft Fabric, including OneLake storage, Bronze/Silver/Gold layers, Delta tables, Fabric Warehouse, and SQL endpoints. Own workspace structure, environment setup, and overall platform health.
Ingestion & Integration: Design pipelines to ingest data from client systems, internal applications, CRMs, REST APIs, SharePoint, Excel, and manual trackers into governed lakehouse tables. Use Fabric Data Factory, Dataflows Gen2, Spark notebooks, OneLake shortcuts, and mirroring to replace manual processes with scheduled, monitored loads.
Data Modeling: Model Silver and Gold layers for analytics, including dimensional models, wide operational KPI tables, and datasets built for Direct Lake semantic models in Power BI. Partner with BI Managers and Analysts to ensure models serve reporting needs upon initial release.
Data Quality & Observability: Build validation checks, reconciliation, and alerting into every pipeline. Define and document correction rules for known source issues (e.g., time and attendance anomalies). Track freshness, completeness, and failures, maintaining clear runbooks for recovery.
Governance, Lineage, & Security: Register datasets, owners, and lineage in the data catalog. Apply role-based access control (RBAC), workspace/item permissions, sensitivity labels, and Entra ID security groups to maintain auditable data access and support client data segregation requirements.
Performance & Cost Optimization: Tune Spark jobs, Delta tables, and SQL to run within Fabric capacity limits. Monitor capacity units (CUs), optimize refresh schedules, and keep capacity costs predictable as data volumes grow.
DevOps & Release Management: Maintain pipelines, notebooks, and models in Git with deployment pipelines across dev, test, and production environments. Automate deployment, inventory, and administration tasks using PowerShell, Python, and Fabric/Power BI REST APIs.
Semantic Layer & AI Support: Collaborate with BI teams on Direct Lake semantic models, refresh strategies, and DAX measure performance. Prepare governed datasets and semantic models for Fabric data agents, Copilot, and natural language access, exposing governed data via APIs and exports for reporting.
Cross-Functional Partnership: Manage the data engineering intake queue with clear estimates and delivery dates. Partner with BI, Performance Management, Workforce Management, IT, Finance, and client teams to align on source access, data availability, and priorities.
Qualifications & Requirements
Experience
3+ years of hands-on data engineering experience building and operating production pipelines, including end-to-end ownership of a data platform.
Core Technical Skills
SQL & Data Modeling: Strong command of SQL, query optimization, window functions, incremental loading patterns, and dimensional modeling (Kimball) for analytics.
Python & PySpark: Proficiency in Python for data engineering, including PySpark notebooks, pandas, and API integrations.
Lakehouse Platforms: Experience with Microsoft Fabric (OneLake, Lakehouse, Delta tables, Fabric Data Factory/Dataflows Gen2, Spark notebooks, Fabric Warehouse) or comparable platforms (Azure Data Factory, Synapse, Databricks).
Architecture: Hands-on experience designing medallion architecture (Bronze, Silver, Gold layers) and navigating structural tradeoffs.
BI & Semantic Layers: Working knowledge of Power BI (semantic models, Direct Lake, Import modes, refresh management, basic DAX for troubleshooting).
Integration & Ingestion: Track record of integrating REST APIs, SaaS applications, CRMs, SharePoint, Excel, and manual sources into automated workflows.
Governance & CI/CD: Familiarity with data quality validation, data governance (RBAC, Entra ID, sensitivity labels, data catalogs), Git-based version control, CI/CD, and scripting via PowerShell or Python.
Soft Skills
Strong written and verbal English communication skills, with a proven ability to document technical work and explain data models to business partners clearly.
Preferred Qualifications
Domain Knowledge: Experience handling operational data from BPO, logistics, contact centers, workforce management, or shared services (e.g., attendance, productivity, service level, and quality metrics).
Certifications: Microsoft Certified: Fabric Data Engineer Associate (DP-600) or Fabric Analytics Engineer Associate.
Tools & Frameworks: Experience with Snowflake, DataHub (or other data catalogs), Fabric data agents/Copilot, Azure DevOps, and Agile/Scrum delivery methodologies.
Mindset & Operational Style
Treats pipelines as products complete with ownership, testing, monitoring, and documentation.
Models data around business questions rather than the raw structure of source files.
Distinguishes between technical data issues and conceptual definition issues, addressing both proactively.
Prioritizes modular, reusable data assets over one-off redundant tables.
Monitors platform cost and capacity limits alongside job runtimes.
Skills
- Agile
- AI
- Analytics
- API
- Automation
- Azure
- Azure Data Factory
- Azure DevOps
- CI/CD
- Data Engineering
- Data Governance
- Data Modeling
- Data Quality
- Databricks
- DevOps
- Dimensional Modeling
- Entra ID
- Git
- Lakehouse
- Microsoft Fabric
- Observability
- pandas
- Performance Management
- Power BI
- PowerShell
- PySpark
- Python
- RBAC
- REST
- SaaS
- Scrum
- SharePoint
- Snowflake
- Spark
- SQL
- Version Control