Data Engineer
Summary
Builds and maintains data pipelines, migrating systems to a new framework while ensuring data quality, integrity, and reliability using Python, Java, Kafka, Airflow, and cloud storage solutions.
- Migrate data pipelines from existing data acquisition framework to the new GDP data acquisition framework
- Configure, develop and deliver data ingestion scripts for loading data into T1 data layer
- Develop and manage ETL/ELT workflows, ensuring data quality, integrity, and reliability.
- Integrate and automate data quality checks and validation processes within data pipelines.
- Deploy and manage containerized applications using Docker and orchestrate workloads on Kubernetes.
- Work with modern data lake and warehouse technologies such as Iceberg
- Orchestrate complex workflows using Airflow.
- Integrate with data catalog and governance tools such as Datahub and Ranger.
- Collaborate with cross-functional teams to understand business requirements and deliver data solutions.
- Ensure security, compliance, and best practices in data management and governance.
Key Skills Required:
- Strong proficiency in Linux, Python, and Shell scripting.
- Hands-on experience with Docker, Kubernetes, and container orchestration.
- Hands-on experience with Minio and Azure Data Lake Storage (ADLS) using S3 protocols.
- Experience with Apache Iceberg, Kafka, Airflow, Datahub, Trino, and Ranger.
- Proficiency in Java for data engineering tasks.
- Solid understanding of data modeling, data warehousing, and big data technologies.