Junior Data Engineer
ABOUT CODEWAY
Codeway builds category-leading consumer AI apps on mobile and web, at global scale.
600M+ downloads. 60+ apps. 400+ builders across Barcelona and İstanbul. Completely bootstrapped. No board. Profitable since year two.
Small teams move faster than big ones. But they need big resources to win at scale. Nobody starts from zero. The problem decides the solution, not us.
We turn category wins into six focused vertical companies: Wishlabs (AI Creativity), Wellture (Wellness), Learna (Education), Sparked (Media), Cosmic Jam (Entertainment), and Catalyst Apps (Utilities & Productivity). Each has full ownership and the depth to go further than anyone else in its space.
Codeway HQ powers them all. One platform. One team that finds and grows builders. One set of guardrails. Every idea born here starts with an unfair advantage: the platform, the people, and the profits of every product before it. That's the system.
Turns out you can do both.
Move like an indie. Hit like a giant.
Position
You'll join Cerebro, the team that builds and runs the data platform behind 60+ Codeway apps and 400M+ users. This is the layer where streaming events become trustworthy numbers — the retention, conversion, and attribution that product, marketing, and the executive team act on every day. You'll start owning real pipelines and warehouse models on a modern Google Cloud stack, with the team alongside you, and grow into more of the platform as you go. You'll also ship with agentic AI tools as a real part of how we work — this is AI-native data engineering. If you like owning data end to end, from a Pub/Sub event to a Looker explore, and want scale you can't get just anywhere (hundreds of millions of users), you'll feel at home here.
What You'll Do
Day-to-day
Build and maintain data pipelines that ingest events reliably into the data platform.
Write SQL and build/maintain warehouse models that feed analytics and BI.
Help keep our BigQuery warehouse fast and cost-aware (partitioning, clustering, sensible schemas).
Contribute to the Looker/LookML semantic layer (PDTs, datagroups, derived views, explores).
Build with agentic coding tools (Claude Code, Google Antigravity) every day, and help keep our internal knowledge base sharp.
Grow into over time
Real-time/streaming ingestion through Pub/Sub and Dataflow.
Transformation and orchestration with dbt and Airflow.
Reading and making changes in our distributed Node.js services (event routing, conversion matching, receipt verification) where pipeline work lands.
Marketing-platform ETLs and attribution/growth KPIs.
You won't touch all of this in your first months — the platform surface spans ingestion, transformation, the warehouse, and the BI layer, and we'll bring you onto it deliberately.
What You'll Bring
An AI-native way of working — you build alongside agentic coding tools (Claude Code, Google Antigravity) and judge their output critically, rather than trusting it blindly.
Solid programming fundamentals — sound programming principles, data structures, and clean, maintainable code. The language matters less than how you think.
Comfortable with SQL and data modeling — you can write non-trivial SQL and reason about how to structure data.
You've built and owned a data pipeline — through a personal project, coursework, an internship, or production work. No prior professional experience required.
Fluency in English — you communicate and document clearly across product, marketing, analytics, and engineering.
A Bachelor's or higher in Computer Science, Software Engineering, or a related field.
Day one we really expect: SQL and one general-purpose language (we work across Node.js/TypeScript and Python). Everything else below, we'll teach.
Nice to have — none is a dealbreaker:
Hands-on with any cloud data warehouse (BigQuery especially) and GCP (Pub/Sub, Dataflow, Cloud Functions, Cloud Run, GKE, Firestore, GCS).
A modern analytics & transformation stack — Looker/LookML, dbt, Airflow.
Comfort with Docker/containers and a feel for how distributed data processing behaves.
A marketing-analytics brain — attribution, retention, conversion, ROAS, LTV.
THE RECRUITING PROCESS
We are committed to keeping our recruitment process short and transparent. Here’s how it looks like:
Application
Case Study
Talent & Culture Interview
Tech Panel Interview
Final Interview
Offer
As published by ashby · 13 questions · 1 written answer
Basics
Name, Email, Location, Resume, Interview Recording Consent
Short answers (4)
- Phone optional
- GitHub optional
- If offered the position, how soon would you be able to start?
Pick from a list (8)
- What is your English proficiency level?
- I confirm that I am aware this role requires relocation to the city Location and I would be open to relocating if the process proceeds positively.
- Will you now or in the future require visa sponsorship to work in Spain?
- How many years of experience do you have?
- Which best describes your SQL experience?
- In which languages have you written non-trivial code — structured into modules and functions, not just notebooks or scripts?
- What's the most autonomous data pipeline you've built and owned end to end?
- What best describes your experience querying datasets with tens of millions of rows or more than 10 GB of data per day?
Written answers (1)
- How do you work with agentic coding tools (Claude Code, Antigravity, Cursor, Copilot)? Please briefly describe your experience: list the tools you used and what you used them for. Please answer in 2 to 4 sentences.