Data Engineer CBT
Summary
Designs and builds data pipelines, warehouses, and APIs to ingest, process, and analyze large datasets using Cloudera, Spark, and Python, then visualizes results in Power BI.
- Build data pipelines ingesting and integrating datasets from multiple data sources, while designing and developing solutions for data integration, data modeling, and data inference and insights.
- Design, monitor, automate, and improve development, test, and production infrastructure for data pipelines and data stores.
- Troubleshoot and performance tune data pipelines and processes for data ingestion, merging, and integration across multiple technologies and architectures including ETL, ELT, API, and SQL.
- Build pipelines and solutions on Cloudera (HDFS) platforms.
- Develop Reports using SQL Server Reporting Services (SSRS) and Dashboards using Power BI.
- Develop Data APIs using Python FAST API framework.
Minimum qualifications:
- Bachelor’s / Master’s degree in Computer or Software engineering
Minimum experience:
- At least 2 years of Experience in dealing with Big Data & Data Warehousing
Required expertise:
- Experience of building data pipelines using any Data Engineering Tool
- Good understanding and hands-on experience of Data Engineering principles, Data warehousing, ETL process, SQL, and handling data in JSON and other semi-structured formats.
- Hands on experience on HDFS, Spark, Impala and KAFKA
- Understanding and proficiency with at least one programming language (Python preferable, Scala, Go)
- Good knowledge of operating systems (Linux power user)
- Quick learner and ability to adapt to customer-driven fast-paced development environments.
- Aptitude to learn new technologies.
- Team player with outstanding collaboration and teamwork attitude.
- Excellent written and verbal communication skills.
- Excellent analytical and problem-solving skills.
- Good to have knowledge on Temenos data model.
- Good to have prior experience developing Reports on Temenos.