Strong Middle Data Engineer
Summary
Designs and builds data pipelines to feed analytics and reporting systems using Python, SQL, and cloud platforms like AWS/GCP.
Jira ticket
We are expanding our team and looking for a Data Engineer to help us build a scalable, reliable data architecture.
In this role, you will build reliable data pipelines, design clear domain models, and optimize our data warehouse. We value a practical, problem-solving mindset — someone who easily navigates ambiguous requirements, brings structure to complex tasks, and takes ownership of their solutions.
In this role, you will:
Build and support Airflow pipelines to ingest data from third-party sources (APIs, payment providers, ad networks) with retries, backfills, and quality checks.
Work directly with Analytics Engineers and product teams to translate ambiguous business requests into reliable data models.
Design and implement fact and dimension tables in dbt using dimensional modeling principles.
Migrate existing SQL transformations into dbt models, adding tests and documentation.
Refactor the current dbt project: improve layer structure, standardize naming conventions, and set up CI checks.
Optimize slow or costly warehouse queries and materializations.
Evaluate table formats and set up initial ingestion flows for our Data Lakehouse prototype.
Skills you’ll need to bring:
Strong proficiency in Python and advanced SQL, including complex transformations, window functions, and query optimization.
Proven experience building and orchestrating ELT/ETL pipelines using Airflow.
Solid experience with dbt for data modeling, testing, documentation, and managing project structure.
Good understanding of dimensional modeling concepts, Kimball methodology, star/snowflake schemas, and SCDs.
Practical experience working with cloud data warehouses (preferably BigQuery, Redshift, or Snowflake) and Cloud Storage.
Hands-on experience with Docker, Git workflows, and basic cloud/containerized environments.
A problem-solving mindset with the ability to handle ambiguous requirements, communicate effectively with stakeholders, and use AI/LLM tools (like Claude) to speed up delivery.
At least an Intermediate level of English and fluent Ukrainian.
As a plus:
Experience with change data capture (CDC) tools and patterns (e.g., Debezium).
Background in building and optimizing large-scale data processing jobs with PySpark.
Experience with real-time data processing using message brokers like Kafka or RabbitMQ.
Experience with stream processing frameworks (Flink, Kafka Streams).