Data Engineer - Automation & Innovation Department
Summary
Build and maintain scalable data pipelines, automate ingestion from Kafka, MQ, SFTP, and databases, and optimize cloud-based processing in GCP using BigQuery, Dataflow, and Airflow.
Type of cooperation: Hybrid ( 2-3 days a week in the office)
Recruitment online!
You will join a diverse, international team of data professionals working across multiple countries. In this role, you will collaborate with both local experts and international colleagues to build a unified, high-performance data infrastructure. We are looking for a Data Engineer who enjoys solving complex integration challenges and wants to have a real impact on how our organization uses data globally.
What tasks await you?
- Build and maintain data ingestion processes from various sources into the Data Lake.
- Design, develop, and optimize complex data pipelines for reliable data flow.
- Build, develop, and maintain frameworks that facilitate the construction of data pipelines.
- Implement end-to-end testing frameworks for data pipelines.
- Collaborate with data analysts and scientists to ensure the delivery of quality data.
- Ensure robust data governance, security, and compliance practices.
- Explore and implement emerging technologies to improve data pipeline performance.
- Utilize and integrate data from various source system types, including Kafka, MQ, SFTP, databases, APIs, and file shares
What skills will be appreciated?
- Minimum 3 years of experience in a Data Engineering role.
- Practical experience with Cloud (preferably GCP)services, including BigQuery, Dataflow, Google Cloud Storage (GCS), Data Catalog, and dbt.
- Advanced SQL proficiency, including writing complex queries, query optimization, Common Table Expressions (CTEs), window functions, and related advanced SQL techniques.
- Python coding skills (experience with Pandas and NumPy is a plus)
- Experience working with Apache Airflow for pipeline orchestration, workflow automation, monitoring, data quality validation, alerting, and troubleshooting.
- Experience with Apache Spark, Hadoop, and Hive for large-scale data processing.
- Experience designing, developing, testing, maintaining, and optimizing ETL/ELT pipelines for batch and streaming data processing, including Apache Kafka, API-based data ingestion.
- Experience designing and implementing data integration solutions, data warehouse architectures, data models, and data layers to support analytics and reporting.
- Experience implementing and maintaining CI/CD pipelines and using GitLab for version control and technical documentation.
- Experience working in Agile/Scrum environments.
- Good command of English.
Nice to have:
- Bachelor’s or Master’s degree in Data Science, Statistics, Computer Science, Economics, or a related field.
- Basic understanding of Machine Learning and MLOps principles
- Strong understanding of data governance principles, including metadata management and data quality frameworks.
- Experience in international or multi-country data projects is preferred.
- Strong communication skills to translate technical findings into business-oriented insights.
- Self-starter with a continuous learning mindset and strong attention to detail.