Data Engineer
Summary
Senior Data Engineer for a leading financial institution in Singapore: builds scalable lakehouse platforms, batch/real-time pipelines, and AI-ready data assets for banking risk and fraud use-cases. Core stack: Spark/PySpark, Java, Python, Docker/Kubernetes, plus RAG/NLP and agent-AI data workloads.
A leading financial institution is seeking a Senior Data Engineer to build scalable lakehouse platforms, governed data products and AI‑ready data assets in close collaboration with architecture, governance, analytics and business stakeholders.
You will deliver batch & real‑time data pipelines, enforce data quality and observability, and partner with AI/data science teams to power NLP, RAG and agent‑driven analytics for banking risk and fraud use‑cases.
Responsibilities:
- Design metadata‑driven ingestion frameworks for structured and unstructured data, covering batch, CDC, streaming, API and file sources
- Develop reusable Spark / Java / Python components; build Bronze‑Silver‑Gold lakehouse layers and implement data modelling, lineage, SLA‑enabled data products
- Tune distributed workloads, implement data quality controls and production observability
- Apply DevOps, CI/CD and IaC on Docker / Kubernetes; operationalise ML models alongside data scientists
- Architect data foundations supporting RAG, agent‑AI and multi‑modal unstructured‑data workflows
- Build internal engineering tooling and align data implementations with financial regulatory requirements
- Troubleshoot production workloads and drive platform reliability improvements
Requirements:
- 10+ years hands‑on data‑engineering delivery experience, banking / financial services background preferred
- Strong skills in PySpark, Java, Python, advanced SQL; solid knowledge of data modelling, lakehouse architecture, ETL/ELT and CDC
- Practical experience building ingestion pipelines and data products; working familiarity supporting RAG / NLP / agent‑AI workloads
- Hands‑on with Docker, Kubernetes, CI/CD and observability tooling
- Able to translate business & risk requirements into data‑platform designs; good cross‑team communication
- Bachelor / Master’s in Computer Science or quantitative‑related discipline
- Experience with Iceberg, Kafka, Flink, NiFi or Cloudera Machine Learning
EA License no.: 16S8066 Rep no.: R25157345
Only successful applicants will be notified.