Data Engineer (ETL, Snowflake, AWS Glue)
Summary
Singapore-based data engineer role centered on building ETL pipelines and migrating legacy Oracle/Hive transformation jobs to Snowflake, AWS Glue and Spark/Python. Day to day involves data ingestion, platform support (BAU), and leading a team of data engineers using SQL, Python, PySpark, Airflow/MWAA and AWS services.
Roles & Responsibilities:
- Develop tools to improve data flows between internal/external systems and the data lake/warehouse.
- Work with stakeholders to understand needs for data structure, availability, scalability, and accessibility.
- Build robust and reproducible data ingest pipelines to collect, clean, harmonize, merge, and consolidate data sources.
- Understanding existing data applications and infrastructure architecture
- Build and support new data feeds for various Data Management layers and Data Lakes
- Evaluate business needs and requirements.
- Support migration of existing data transformation jobs in Oracle, and MS-SQL to Snowflake.
- Lead the migration of the existing data transformation jobs in Oracle, Hive, Impala etc. into Spark, Python on Glue etc.
- Able to document the processes and steps.
- Develop and maintain datasets.
- Improve data quality and efficiency.
- Lead Business requirements and deliver accordingly.
- Collaborate with Data Scientists, Architect and Team on several Data Analytics projects.
- Collaborate with DevOps Engineer to improve system deployment and monitoring process.
- Experience in critical production support and how the BAU functions is preferred.
Key Requirements:
- Lead a team of data engineers in managing day-to-day BAU operations, production support, incident management, and platform stability.
- Strong AWS knowledge in terms of designing new architecture and providing optimized solutions for existing ones. (S3, Glue, DMS, MWAA, AMS, IAM).
- In-depth knowledge with respect to Snowflake and its architecture.
- Prefer prior Experience in BAU environment.
- Good knowledge on Airflow and MWAA.
- Hands-on experience in SQL/Python/Pyspark.
- Expertise in optimizing techniques in cloud environments.
- Should have the vision on data strategy and able to deliver the same.
Must Have: AWS - Glue, Airflow, Lambda, S3 , Snowflake (Related Cloud DB) SQL ,Python/ Spark,CI/CD Good To have : ECS,GIT, Oracle