Lead Data Engineer
Summary
Lead Data Engineer (hybrid, Pune) building real-time and batch data pipelines for AI-based supply chain solutions — designing ETL/ELT workflows, integrations, and data quality/monitoring controls with Python, PySpark, advanced SQL, MongoDB, Docker and Kubernetes, plus exposure to agentic AI designs.
🚀 We’re Hiring: Lead Data Engineer – Supply Chain
Role: Lead Data Engineer
Domain: Data Engineering | Supply Chain | AI
Experience: 6+ Years
Location: Pune
Work Mode: Hybrid
About the Role
We are looking for an experienced Lead Data Engineer with strong hands-on expertise in Data Engineering and Supply Chain to design and build scalable, real-time, and batch data platforms.
The ideal candidate should have strong experience in Python, PySpark, Advanced SQL, data transformations, system integrations, and Supply Chain workflows, along with exposure to AI and Agentic solutions.
🛠️ Key Responsibilities
▪️ Design and build real-time and batch data ingestion pipelines for AI-based Supply Chain solutions.
▪️ Develop complex data transformation, cleansing, enrichment, and reconciliation workflows.
▪️ Integrate data from REST APIs, databases, files, enterprise systems, and streaming sources.
▪️ Develop ETL tools and validation routines for complex Supply Chain workflows.
▪️ Build reusable Python and SQL components for ingestion and transformation.
▪️ Implement schema evolution, incremental loads, retries, error handling, and data recovery.
▪️ Understand agentic structures and create data solutions/models for optimization engines.
▪️ Implement data quality checks, monitoring, logging, alerting, and audit controls.
▪️ Review technical designs and guide data engineers on implementation standards.
Required Skills
✅ 6+ years of hands-on experience in Data Engineering
✅ Strong experience with Supply Chain data management and transformation
✅ Strong Python & PySpark development skills
✅ Advanced SQL, including complex joins and query optimization
✅ Strong understanding of ETL/ELT design patterns
✅ Experience with REST APIs, JSON, CSV, MongoDB, and OLAP engines
✅ Experience developing Agents and Agentic solutions/designs
✅ Strong understanding of data modelling, schema design, partitioning, and schema evolution
✅ Experience with pipeline observability, data validation, monitoring, and alerting
✅ Hands-on experience with Git, CI/CD, Docker, and Kubernetes
Interested candidates can share their updated resume at: prajakta@prosapiens.in
WhatsApp: 8055008782
🤝 References would be highly appreciated!