Senior Data Engineer (Banking, 1-year renewable contract)
EVOLUTION RECRUITMENT SOLUTIONS PTE. LTD. Senior Data Engineer (Banking, 1-year renewable contract)
Dear Applicant,
If you or someone you know is interested, please send the CV directly to quynh.nguyen@evolutionjobs.sg (most preferred, as I may overlook some CVs due to the high volume).
Please note that visa sponsorship is not available at this time.
Key Responsibilities
- Implement and operationalize enterprise-scale Lakehouse platforms, data products, and data marketplace capabilities.
- Design, develop, test, and maintain scalable batch, streaming, CDC, and API-based data ingestion pipelines.
- Develop multimodal data ingestion pipelines, including:
o Content extraction from various file formats.
o Regex-based extraction of specific fields.
o Content extraction from embedded images.
o Frame extraction from video files.
o Transcript extraction from audio files.
- Build and maintain foundation and business data products with defined data contracts, SLAs, data quality controls, and governance standards.
- Implement and work with modern open table formats, including Iceberg, Hudi, and Delta Lake.
- Develop and support data pipelines for RAG, vector search, Generative AI (GenAI), and agentic AI use cases.
- Perform performance tuning, optimization, production support, troubleshooting, and root cause analysis across data platforms and pipelines.
- Implement data ingestion, transformation, reconciliation, and data quality frameworks.
- Develop data architectures supporting NLP and AI-driven analytics, including the ingestion, curation, governance, and management of structured and unstructured data.
- Support ML platforms and workflows, including model development, deployment, and operationalization.
- Develop internal engineering tools and full-stack applications using Python, shell scripting, and modern web frameworks.
- Expose and integrate data through APIs, event streams, dashboards, and BI platforms.
- Create and maintain technical documentation, deployment guides, and operational runbooks.
- Ensure compliance with engineering standards, DevSecOps controls, CI/CD practices, security requirements, and software delivery standards.
- Collaborate effectively with distributed engineering, data, technology, and business teams across multiple projects.
- Drive continuous improvement in data platform reliability, scalability, automation, and operational excellence.
Key Requirements
- Bachelor’s degree in Computer Science, Engineering, Information Technology, or a related discipline.
- 8–12 years of experience in Data Engineering, Big Data, Data Lake, Lakehouse, or large-scale data platform implementations.
- Strong hands-on experience with enterprise data platforms such as Databricks, Snowflake, Cloudera, Azure, AWS, GCP, Huawei, or Alibaba Cloud.
- Hands-on experience designing and developing data products and data marketplace capabilities.
- Strong expertise in Spark, PySpark, SQL, Python, and/or Scala, with strong overall programming skills.
- Proven experience building scalable data ingestion, transformation, streaming, CDC, reconciliation, and data quality pipelines.
- Hands-on experience with Iceberg, Hudi, Delta Lake, and object storage platforms.
- Experience with modern data technologies such as Kafka, Flink, Spark Streaming, Airflow, Trino, Dremio, Hive, and/or Impala.
- Strong hands-on experience with Kubernetes, OpenShift, Docker, Terraform, Jenkins, Git, CI/CD, MLflow, and observability tools.
- Experience designing data architectures and pipelines supporting NLP, AI-driven analytics, RAG, vector search, GenAI, and agentic AI use cases.
- Experience handling and governing unstructured and multimodal data, including documents, images, audio, and video.
- Experience with ML platforms and frameworks such as CML, Spark MLlib, scikit-learn, and XGBoost, including model deployment.
- Strong knowledge of data modelling, metadata management, data lineage, data governance, and data quality.
- Experience exposing data through APIs, event streams, dashboards, and BI platforms.
- Experience with full-stack/internal engineering tool development using Python, shell scripting, Flask, React, or similar technologies is advantageous.
- Knowledge or experience with Teradata, Netezza, Greenplum, or MPP migration programmes is advantageous.
- Strong engineering, automation, troubleshooting, and performance optimization mindset.
- Strong communication and stakeholder management skills, with the ability to work effectively across distributed teams and multiple projects.
- Experience working in Agile delivery environments and enterprise-scale technology platforms.
- Strong commitment to quality, operational excellence, automation, and continuous improvement.
- Relevant certifications such as Databricks Certified Data Engineer, Azure Data Engineer Associate, AWS Data Analytics Specialty, Google Professional Data Engineer, SnowPro, or DAMA CDMP would be an advantage.
Skills
- Agentic AI
- Agile
- AI
- Airflow
- Analytics
- API
- Automation
- AWS
- Azure
- Bash
- CI/CD
- Cloud
- Data Analytics
- Data Engineering
- Data Governance
- Data Ingestion
- Data Lake
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Quality
- Databricks
- Delta Lake
- DevSecOps
- Docker
- Flask
- Flink
- GCP
- Generative AI
- Git
- Hive
- Jenkins
- Kafka
- Kubernetes
- Lakehouse
- Machine Learning
- Metadata Management
- MLflow
- Model Deployment
- NLP
- Observability
- OpenShift
- PySpark
- Python
- React
- Scala
- scikit-learn
- Snowflake
- Spark
- SQL
- Stakeholder Management
- Teradata
- Terraform
- Trino
- Vector Search
- XGBoost