Data & Integration Engineer
Summary
Designs and builds data integrations between systems using APIs, SFTP/file transfers and batch pipelines within a DataLake ecosystem (Informatica, Cloudera), and prepares structured/unstructured data for GenAI workflows like RAG, while coordinating deliveries across data, application, infrastructure and security teams.
1. System Analysis & Design
Analyse business/technical requirements and translate them into
data flows and integration designs
Work with upstream and downstream teams to define
data contracts and interfaces
Identify gaps, inefficiencies and risks in current data movement processes
Propose pragmatic solutions balancing speed, quality and maintainability
2. Integration & Data Movement
Design and implement
data movement across systems
using:
APIs
SFTP and file based transfers
Batch pipelines
Coordinate integrations across systems in the DataLake ecosystem
(Informatica, Cloudera, etc.)
Ensure data is correctly transformed, mapped and delivered to target systems
Troubleshoot integration issues across environments
3. Data Preparation for GenAI
Support data ingestion and preparation for GenAI use cases:
document ingestion
data aggregation
enrichment and transformation
Work with structured and unstructured data
Ensure data is usable for downstream AI workflows (RAG, search, investigation flows)
You are not asking them to build models, just make data usable for them.
4. Delivery & Coordination
Work across multiple teams:
data platforms
application teams
infrastructure
security
Support SIT, UAT and production rollouts
Ensure integration reliability, error handling and monitoring
Document flows, mappings and interfaces clearly