Point your AI agent at freehire and let it find you a job.

Get the CLI →

Cato

NewBe an early applicant

Data engineer

Posted
Discussion

Summary

Own end-to-end data pipelines at Cato that scrape national tender portals in Spain and Italy, then merge, deduplicate, enrich (LLM extraction, embeddings, OCR) and deliver tenders to customers. Core stack: Python, PostgreSQL, Prefect, K3s/ArgoCD, AWS; on-site in Barcelona.

Your mission


Build and run the pipelines that carry a tender from the source portal to the customer's screen: ingestion, merge, enrichment, delivery. You own concrete pieces of the data pipeline end to end — not tickets handed to you, but the sources, jobs and tables behind them.


What you'll actually do

  • Own scrapers and ingestion for a set of national portals across Spain and Italy, where reading the source in its original language is part of the job.
  • Write and maintain orchestrator flows: retries, backfills, alerting, and a clear answer to \"Did today's run actually land?\"
  • Work on merge, dedup and reconciliation — the same tender arrives three times, in three shapes, and only one version can reach the customer.
  • Ship AI enrichment steps: batch LLM extraction of requirements, embeddings, OCR on attachments.
  • Write SQL that survives production: query plans, indexes, JSONB, partitioning, CONCURRENTLY migrations.
  • Guard data quality with tests and checks that fail loudly before a customer finds the gap.


Ideal profile

  • Real SQL: you can read an EXPLAIN and say why the plan is bad, not just that it is slow.
  • Python you'd put in production: typed, tested, and readable six months later.
  • Pipelines you've actually operated: with an orchestrator (Prefect, Airflow, Dagster, ArgoWorkflows) and the 3 a.m. failures that come with them.
  • Builder by default: you see a manual process and your first instinct is to automate it.
  • Comfortable with messy sources: broken HTML, inconsistent XML, PDFs that were scans of scans.


Experience

  • 2–4 years building data pipelines in production.
  • Hands-on with PostgreSQL beyond writing queries — you've had to make one fast.
  • Exposure to LLM-based extraction is welcome; curiosity about it is mandatory.


What you won't find here

  • No micromanagement: we trust you to own your part of the stack.
  • No \"standard\" 9-to-5 mentality: we care about outcomes and we are looking for people who are willing to go the extra mile.
  • No \"we've always done it this way\" excuses: we're here to disrupt, not to follow old patterns.


Our Tech Stack

  • Data & Infra: Python, PostgreSQL, Prefect, K3s/ArgoCD, AWS
  • AI: batch LLM extraction, embeddings, OCR


Compensation

RAL €40,000 – €60,000 + equity, depending on seniority and profile.


Hiring Manager

Lorenzo Rossetto

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available