Point your AI agent at freehire and let it find you a job.

Get the CLI →

Deep Knowledge Group

NewBe an early applicant

Data Engineer — Data Collection and Enrichment

Discussion

Summary

Build and own data collection and enrichment pipelines using Python and SQL/Postgres, leveraging LLMs for extraction and classification to maintain clean, verifiable intelligence databases.

About the work

We build and maintain industry intelligence databases - organizations, people, products and their relationships across sectors such as HealthTech, QuantumTech, LegalTech and wellness, covering multiple regions. Every one of them starts as messy public data and has to end up as a clean, deduplicated, verifiable dataset that analysts and downstream products can rely on.

That means the hard problems here are not "write a scraper". They are: how do you know two records are the same organization? How do you prove where a field came from six months later? How do you re-run a collection over a source that changed its markup, without corrupting what you already have?

You will work directly with the Head of Data Science and a team of AI and data engineers.

What you will do

Build and own collection pipelines over public web sources - structured, semi-structured and awkward
Design the enrichment layer: normalization, deduplication, entity resolution, confidence scoring
Keep provenance on every field, so any value can be traced back to its source and date
Use LLMs where they genuinely help - extraction from unstructured text, classification, taxonomy assignment - and know where they don't
Take datasets from one-off collection to scheduled, monitored, re-runnable pipelines
Define and enforce quality checks: coverage, freshness, duplicate rate, field completeness

What we're looking for

Both levels

Strong Python - you write services and pipelines, not notebooks
Real production scraping experience: sessions, rate limits, retries, blocks, anti-bot, JS-rendered pages, APIs that are not documented
SQL and PostgreSQL beyond CRUD - indexing, query plans, schema design
Practical experience with LLM APIs for extraction and classification, including how you validate their output
Data quality instinct: you check what you collected before you hand it over

Additionally for the Senior opening

You have owned a data platform, not just tasks on one - orchestration, scheduling, monitoring, backfills
Experience with entity resolution or record linkage across sources that disagree with each other
Airflow, Dagster or Prefect in production
You have set the standards other people then worked to: schemas, conventions, review

Nice to have

Docker and CI/CD
Knowledge graphs, ontologies, or taxonomy design
Vector search and RAG pipelines
OSINT methodology and source verification
Experience with data licensing or compliance constraints on public-web collection

How we hire

Intro call - 30 minutes, mutual fit and what you've actually built
Paid proof task - a real, bounded piece of our work. Paid at market rate, roughly 4 hours. We are not asking anyone to work for free
Technical review - we go through your solution together: your trade-offs, what you'd do with more time, what you'd do differently at ten times the volume
Offer - band determined by the level you land at, not by the title you applied under

We evaluate the proof task on judgment, not polish. Reproducibility, honest handling of edge cases, and clear reasoning about what you chose not to do count for more than a clever one-liner.
What we care about

Two things, mostly.

You question your own data. The engineers who work out well here are the ones who notice that a source silently changed its schema, that a "97% match rate" is hiding a systematic bias, that a field is populated but wrong. Volume without verification is worth nothing to us.

You explain your reasoning. Much of this work involves judgment calls that nobody can check quickly - which record wins a merge, what confidence threshold to use, when to stop enriching. We need those decisions written down and defensible, not buried in a script.

To apply: send your CV plus one short paragraph on the hardest data collection or deduplication problem you have solved, and what you got wrong on the way.

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available