Data Engineer z GCP (f/m/x)
Summary
Build and maintain scalable data pipelines on Google Cloud Platform: BigQuery modeling and cost optimization, workflow orchestration with Apache Airflow/Cloud Composer, Terraform-based infrastructure as code, and CI/CD for data solutions, while collaborating with analytics, BI, and product teams. Requires 4+ years of data engineering experience, strong SQL and Python, and fluency in Polish.
This translation is generic and may include errors. Translate into Ukrainian.
Apache Airflow / Cloud Composer Terraform Data Built Tool Apache Spark Databricks Snowflake Microsoft Fabric
Do you want to develop in cloud technologies and work with real data? Join our Data & Analytics team, where we build and develop solutions based on GCP. Work with experts, grow towards Data Engineering, Big Data, or Machine Learning, and have a real impact on projects.
- Design, implement and maintain scalable data pipelines based on Google Cloud Platform
- Work with BigQuery as the main data warehouse: data modeling, query and cost optimization, ensuring performance and reliability of solutions
- Integrate data from various sources (files, databases, APIs, events) and process and transform it
- Orchestrate data workflows using Apache Airflow / Cloud Composer
- Create and maintain CI/CD solutions for data pipelines and infrastructure
- Manage cloud infrastructure according to the Infrastructure as Code approach (Terraform)
- Ensure data quality, monitor pipelines, and respond quickly to incidents
- Collaborate with analytics, BI, and product teams to deliver stable and well-documented data
- Participate in the development of data architecture and jointly define best practices in data engineering
Requirements
- Min. 4 years of experience in a Data Engineer role or similar position working with data in a production environment
- Strong knowledge of Google Cloud Platform, particularly: BigQuery (data modeling, query optimization) and Cloud Storage
- Ability to design, build, and maintain data pipelines (batch and/or streaming)
- Strong knowledge of SQL and Python in the context of data processing and orchestration
- Experience in workflow orchestration (Apache Airflow / Cloud Composer)
- Experience implementing CI/CD for data solutions, e.g., GitHub Actions, GitLab CI, Cloud Build
- Familiarity with the Infrastructure as Code approach, particularly Terraform
- Previous work with large data volumes, considering performance and reliability of solutions
- Required to be located in Poland and fluent in Polish
Nice to have
- Practical experience in processing streaming data (e.g., Dataflow / Apache Beam, Pub/Sub)
- Proficiency in Apache Spark / PySpark when working with large data volumes
- Skills in data transformation and modeling using tools like dbt
- Ability to work with various data platforms (e.g., Databricks, Snowflake, MS Fabric)
- Familiarity with tools and best practices in Data Governance, Data Lineage, and Data Quality
Job no. JOB-2AYDA
Sii ensures that all hiring decisions are made solely on the basis of qualifications and competence. We are committed to equal and fair treatment of all, regardless of legally protected characteristics. At Sii, we promote a diverse and inclusive work environment, in full compliance with applicable anti-discrimination laws.
Remote Hybrid Office
Skills
- Airflow
- Analytics
- API
- BigQuery
- CI/CD
- Cloud
- Data Engineering
- Data Governance
- Data Lineage
- Data Modeling
- Data Pipelines
- Data Quality
- Data Warehousing
- Databricks
- dbt
- GCP
- GitHub
- GitHub Actions
- GitLab
- Infrastructure as Code
- Machine Learning
- Microsoft Fabric
- PySpark
- Python
- Snowflake
- Spark
- SQL
- Terraform
- Workflow Orchestration