Data Engineer
You will be working with cutting edge cloud technologies (GCP, AWS, BigQuery, Databricks, K8s) and building a large scale data infrastructure for analytics, machine learning, and streaming/CDC data delivery.
- Build and operate batch and streaming ingestion into a layered BigQuery DWH (raw → ODS → data marts) using Airflow, Debezium CDC over Kafka with protobuf, Pub/Sub, and Dataflow
- Integrate external data sources end-to-end — marketing platforms (GA4, AppsFlyer, TikTok/Meta/Google Ads), payment providers, S3 buckets, and third-party APIs — including schema contracts, backfills, and reconciliation
- Engineer the data platform itself in Python: custom Airflow operators and connectors in a shared ETL framework, Kafka Connect on Strimzi (K8s), Cloud Functions, and API integrations with external providers
- Build CI/CD and change-management tooling for BigQuery: GitHub-based test-and-approval flows, SQL migration engines (Liquibase/Flyway/Bytebase), sandbox validation, backup and rollback
- Own reliability and correctness of pipelines: idempotency, deduplication, late-data handling, backfill and replay, freshness monitoring and alerting; write integration and unit tests
- Drive data governance and compliance: ITGC-compliant change management for BigQuery, IAM and least-privilege access, PII policy tags and DLP, Unity Catalog on Databricks, column-level lineage (OpenMetadata/Dataplex), and disaster-recovery planning
- Build internal data tools and platform services for agentic workflows with data — Streamlit apps, Slack bots, LLM-based agents and MCP servers that help teams find and use data
- Support analysts and business teams with data requests, fostering data-driven decision-making across the company
- Contribute to system design and architecture with the development team
- Strong practical Python: clean, well-structured, and tested code for services, tooling, and data pipelines
- Solid software design skills (OOP, modularity, design patterns) — we build platform tools for agentic workflows with data and plan to develop data-related backend services, so well-designed code is highly valued
- Experience building and operating services in a cloud environment (GCP, AWS or similar): CI/CD, containerization, monitoring and alerting
- Familiarity with Kubernetes and Terraform — our infrastructure runs on GCP/K8s
- Hands-on experience with DWH-related tasks (BigQuery or another cloud warehouse) and confident working SQL
- Clear communication with non-engineering stakeholders — a meaningful share of the work is data requests from analysts and business teams
- Demonstrated ability to take ownership of technologies or services and proactively contribute ideas to the team
- Advanced SQL: complex queries, window functions, partitioning, clustering, and cost optimization
- Experience building reliable pipelines around CDC (e.g., Debezium): idempotency, schema evolution, backfills, and reconciliation
- Analytical data modeling skills: table grain, facts vs dimensions, slowly changing dimensions, and metric definitions
- Experience with stream processing frameworks such as Flink or Apache Beam/Dataflow
- Exposure to data governance and audit compliance (ITGC/SOX), Databricks Unity Catalog, or lineage/catalog tooling (OpenMetadata, Dataplex)
- Interest in building LLM-based agents and AI tooling for data
- Help us challenge injustice by creating fair choices for millions of people across 48 countries.
- Develop your professional skills with access to mentoring, career consulting, and learning programs.
- Collaborate with teams around the world and gain international experience through our Global Talent Exchange Program.
- Engage in company-wide challenges, awards, sports activities, employee-led social impact and volunteering projects.
- Work alongside people who take initiative, speak openly, and challenge themselves to grow.
- Improve your language skills through co-financed courses and internal speaking clubs.