Data Engineer
Summary
Data Engineer building ETL/ELT pipelines, cleaning and transforming data, and designing databases using SQL, Python, AWS cloud services, and big data frameworks like Spark and Kafka.
- Proficient in general data cleaning and transformation (e.g. SQL, pandas, R, etc) to ensure data accuracy and consistency.
- Proficient in building ETL pipeline (eg. SQL Server Integration Services (SSIS), AWS Database Migration Services (DMS), Python, AWS Lambda, ECS Container task, Event bridge, AWS Glue ,Spring).
- Proficient in database design and various databases (e.g. SQL, PostgreSQL, AWS S3, Athena, mongo db, postgres/gis, mysql, sqlite, volt db, cassandra, etc).
- Experience in cloud technologies such as GPC, GCC (i.e. AWS, Azure, Google Cloud).
- Experience and passion for data engineering in a big data environment using Cloud platforms such as GPC, GCC(i.e. AWS, Azure, Google Cloud).
- Experience with building production-grade data pipelines, ETL/ELT data integration.
- Knowledge about system design, data structure and algorithms.
- Familiar with data modelling, data access, and data storage infrastructure like Data Mart, Data Lake, Data Virtualisation and Data Warehouse for efficient storage and retrieval.
- Familiar with rest api and web requests/protocols in general.
- Familiar with big data frame works and tools (eg. Hadoop, Spark, Kafka, RabbitMQ).
- Familiar with W3C Document Object Model and customized web scraping (e.g. Beautiful Soup, CasperJS, PhantomJS, Selenium, Nodejs, etc).
- Familiar with data governance policies, access control and security best practices.
- Comfortable in at least one scripting language (eg. SQL, Python).
- Comfortable in both windows and linux development environments.
- Interest in being the bridge between engineering and analytics.