Data Engineer (AVP)
Summary
Data Engineer (AVP) at OCBC Bank in Singapore builds, maintains, and improves batch and streaming data pipelines that feed enterprise data warehouse/lakehouse platforms and AI knowledge bases (RAG). Core stack includes Spark, SQL, Python, Airflow, Kafka/Flink, GCP/AWS with Docker/Kubernetes, and Terraform.
- Build and maintain data pipelines feeding data warehouse / lakehouse platforms (e.g., Cloudera, AWS Redshift or GCP BigQuery)
- Implement data models to support reusable analytical datasets and reporting foundations
- Follow established SLAs and monitoring practices, and help troubleshoot pipeline issues
- Develop and maintain batch data processing jobs using Spark, SQL, Python or Java
- Support and contribute to real‑time streaming pipelines using Flink or similar tools, under senior guidance
- Assist in building and operating CDC pipelines (e.g., Debezium, Confluent or Fivetran)
- Build and maintain ingestion pipelines from APIs, GA4, and other data sources
- Implement messaging/streaming integrations using Pub/Sub and Kafka
- Write clean, well‑tested SQL and Python scripts and data ingestion and processing pipelines
- Develop and maintain Airflow DAGs for scheduled and event‑driven workflows
- Follow orchestration best practices established by senior engineers
- Deploy and support data workloads on Cloudera, GCP (Docker, Kubernetes, Cloud Run), AWS equivalent
- Contribute to and maintain CI/CD pipelines
- Use Terraform to provision and manage infrastructure under senior guidance
- Build and support REST APIs and backend services using Python / Flask
- Use Redis caching to meet performance requirements for data services
- Assist in building and maintaining vector database pipelines and embedding generation jobs under senior guidance
- Support to deliver processing and chunking workflows that feed AI knowledge bases and RAG pipelines
- Build, test and monitor semantic search / retrieval quality for AI‑facing data layers
- Work closely with senior data engineers, AI teams, and business stakeholders (Risk, Marketing, Operations)
- Help translate business requirements into technical implementation tasks
- Participate in code reviews and contribute to a culture of continuous improvement
- Diploma, Bachelor's or Master's degree in computer science or a related field
- At least 5 years of experience in data engineering, data platforms, or related roles.
- Solid hands-on experience in build data pipelines on modern data architectures including Data Warehouse, Lakehouse, and batch/real-time data processing systems.
- Working AI knowledge base concepts: vector databases, embeddings, and semantic search is a plus
- Hands-on RAG pipeline components such as document chunking, embedding generation, and retrieval is a plus
- Good experience on building data layers that support LLM / AI agent use cases is a plus.
- Experience in banking or financial services is a plus
- Data Warehouse / Platform: exposure to Cloudera, BigQuery, Redshift, or similar
- Batch Processing: Spark, SQL, ETL, Python, Map/Reduce
- Streaming: familiarity with Flink or other real‑time processing engines (Good to have)
- CDC: exposure to Debezium, Confluent, Fivetran, or similar (Good to have)
- Data Ingestion: APIs, GA4, Pub/Sub, Kafka, Python pipelines
- Orchestration: Airflow or equivalents
- Cloud & Infrastructure: GCP or AWS; basic Docker/Kubernetes/Cloud Run experience
- DevOps / DataOps: exposure to CI/CD pipelines and Terraform
- Backend & Serving: Python, Flask, REST APIs; familiarity with Redis a plus
- Eagerness to learn real‑time / streaming architectures and low‑latency system design
- Basic exposure to LLM applications, RAG, or AI agent concepts is a plus, not required
- Good product mindset and willingness to treat data as a product, not just a project
- Strong communication skills and comfort collaborating with both technical and business stakeholders
- Competitive base salary and comprehensive benefits.
- Strong learning and development opportunities.
- Exposure to impactful, enterprise‑scale Data and AI initiatives across the OCBC Group.
- A collaborative environment that values innovation, craftsmanship, and continuous improvement.
Skills
- AI
- Airflow
- API
- Automation
- AWS
- BigQuery
- CI/CD
- Cloud
- Data Engineering
- Data Ingestion
- Data Pipelines
- Data Warehousing
- Debezium
- DevOps
- Docker
- Embeddings
- ETL
- Fivetran
- Flask
- Flink
- GCP
- Google Analytics
- Java
- Kafka
- Kubernetes
- Lakehouse
- LLM
- Python
- Redis
- Redshift
- REST
- Semantic Search
- Spark
- SQL
- Terraform
- Vector Databases