Lead Data Engineer
Summary
Lead Data Engineer in Bangalore owning enterprise data architecture and building scalable ETL/ELT pipelines for US healthcare data (EHR, RCM, claims), supporting BI, analytics, and AI/ML workloads. Core stack includes SQL, Python, relational and analytical databases, workflow orchestration (e.g., Apache NiFi), and cloud platforms (AWS/Azure/GCP).
Responsibilities
- Own and evolve the enterprise data architecture supporting operational, analytical, and AI workloads
- Design scalable data platforms for transactional databases, analytical data warehouses/lakehouses, and AI/ML data pipelines
- Define data models, integration standards, metadata management, and data lifecycle strategies
- Establish best practices for data engineering, architecture, performance optimization, scalability, reliability, and maintainability
- Evaluate and recommend emerging technologies and architectural improvements
- Design, develop, and optimize robust ETL/ELT pipelines for structured and unstructured data
- Build reliable batch and real-time data integration pipelines from EHRs, Practice Management Systems, APIs, flat files, and third-party healthcare applications
- Develop and optimize workflows using tools such as Apache NiFi or equivalent orchestration platforms
- Ensure high data quality, integrity, consistency, lineage, and observability across all data platforms
- Support relational, NoSQL, and distributed data platforms
- Design and maintain data platforms supporting Business Intelligence, advanced analytics, and machine learning workloads
- Build data pipelines that enable AI/ML model training, feature engineering, vector databases, Retrieval-Augmented Generation (RAG), and LLM/SLM applications
- Collaborate with Data Scientists and AI Engineers to operationalize ML models and AI solutions
- Support MLOps and data versioning best practices
.
Qualifications/Criteria
- Bachelor’s degree in computer science, Software Engineering, or a related field.
- 10+ years of experience in Data Engineering, Data Platform Engineering, or Data Architecture.
- Minimum 5 years of experience working with US Healthcare data, preferably Revenue Cycle Management (RCM), Claims, EHR, or Healthcare Analytics.
- Equivalent practical experience with demonstrated technical leadership will also be considered.
- Proven experience designing enterprise-scale data architecture.
- Strong expertise in SQL and data modeling.
- Hands-on experience with relational databases (PostgreSQL, SQL Server, MySQL, Oracle) and analytical databases/warehouses.
- Experience building scalable ETL/ELT pipelines and workflow orchestration.
- Strong knowledge of batch and streaming data processing.
- Experience with Python for data engineering and automation.
- Experience designing cloud-based data platforms (AWS, Azure, or GCP).
- Working knowledge of modern data lake house architectures.
- Understanding of AI/ML data engineering concepts, including feature stores, vector databases, embeddings, LLMs, and SLMs.
- Strong understanding of data governance, metadata management, data quality, security, and access control.
- Excellent problem-solving, communication, and stakeholder management skills.
Skills
- AI
- Analytics
- API
- Automation
- AWS
- Azure
- Cloud
- Data Engineering
- Data Governance
- Data Lake
- Data Modeling
- Data Pipelines
- Data Quality
- ELT
- Embeddings
- ETL
- Feature Engineering
- GCP
- LLM
- Machine Learning
- Metadata Management
- MLOps
- MySQL
- NiFi
- NoSQL
- Observability
- Oracle
- PostgreSQL
- Python
- RAG
- SQL
- SQL Server
- Stakeholder Management
- Vector Databases
- Workflow Orchestration