Lead Data & Integration Engineer
Responsibilities:
1. System Analysis & Design
- Analyse business/technical requirements and translate them into data flows and integration designs
- Work with upstream and downstream teams to define data contracts and interfaces
- Identify gaps, inefficiencies and risks in current data movement processes
- Propose pragmatic solutions balancing speed, quality and maintainability
2. Integration & Data Movement
- Design and implement data movement across systems using:
- APIs
- SFTP and file based transfers
- Batch pipelines
- Coordinate integrations across systems in the Data Lake ecosystem (Informatica, Cloudera, etc.)
- Ensure data is correctly transformed, mapped and delivered to target systems
- Troubleshoot integration issues across environments
3. Data Preparation for GenAI
- Support data ingestion and preparation for GenAI use cases:
- document ingestion
- data aggregation
- enrichment and transformation
- Work with structured and unstructured data
- Ensure data is usable for downstream AI workflows (RAG, search, investigation flows)
4. Delivery & Coordination
- Work across multiple teams:
- data platforms
- application teams
- infrastructure
- security
- Support SIT, UAT and production rollouts
- Ensure integration reliability, error handling and monitoring
- Document flows, mappings and interfaces clearly
Requirements
Below are the key skillsets that will be required for all relevant tasks mentioned:
· 10 years of experience in system analysis, integration engineering, data engineering or technical delivery roles.
· Strong ability to translate requirements into system flows, data flows, interface specifications and implementation plans.
· Experience working with upstream and downstream teams to define and deliver enterprise integrations.
· Practical experience with REST APIs, SFTP, batch processing, file based integration and data pipeline orchestration.
· Good understanding of data mapping, transformation, aggregation, reconciliation and data quality controls.
· Good SQL skills and basic to moderate Python skills for data handling, scripting, automation and troubleshooting.
· Exposure to Java
· Exposure to Informatica, Cloudera or similar enterprise data platforms.
· Working knowledge of Git, branching, pull requests, code reviews and controlled release practices.
· Familiarity with CI/CD, Jira, Confluence and enterprise deployment processes.
· Experience with Control M or equivalent scheduling tools.
· Familiarity with logging (OTEL) and monitoring tools such as Splunk Elastic Stack.
· Exposure to GenAI concepts such as document ingestion, RAG, embeddings and data preparation for AI workflows.
Key Domain/ Technical Skills:
· Data Engineering,
· System Integrations,
· Python, SQL