Point your AI agent at freehire and let it find you a job.

Get the CLI →

Smirnov Labs

New

Senior Data Engineer

Posted 3 views
Discussion

Summary

Fully remote senior data engineer, embedded (outstaff) with a client that runs a custom AI system: build and own third-party API connectors and end-to-end data pipelines — ingestion, normalization, embeddings/vector index freshness, data quality. Core stack: Python, SQL, Postgres, a workflow orchestrator (Airflow/Dagster/etc.), Docker/CI/CD on AWS or GCP.

## About Us
Smirnov Labs is a fast-growing engineering company founded by an ex-Googler. We help product teams launch and scale quickly by leveraging modern technology and AI-assisted development.
Our projects are diverse — we partner with clients at the earliest stages and own the technical delivery end-to-end.
We're a small, senior-heavy team where every engineer has real impact. No bureaucracy, no hand-holding — just sharp people shipping fast.
This is an outstaff role: you'll be embedded directly with a client's engineering team as part of our delivery group.

## The Role
The client runs a custom AI system, and it's only as good as the data reaching it. Your job is to get that data in: connect to whatever third-party API holds it, pull it reliably, shape it, and keep it flowing.
That means a lot of integration work — every source has its own auth, its own pagination, its own rate limits, and its own idea of what a schema is. You'll own those connectors and the pipelines behind them end-to-end.

## What You'll Do

  • Build and own integrations with third-party APIs — REST, GraphQL, webhooks, the occasional CSV drop or legacy SOAP endpoint. Auth flows, pagination, rate limits, retries, backfills.
  • Design and run the pipelines that feed the client's AI system: ingestion, normalisation, enrichment, and delivery into the stores it reads from.
  • Build the ingestion path for AI workloads — chunking, embeddings, and keeping vector indexes fresh as source data changes, without full re-indexing every time.
  • Model the data. Design schemas that hold up as sources are added and change underneath you.
  • Make pipelines idempotent, observable, and recoverable. Things will break upstream; the system should degrade predictably and tell you why.
  • Own data quality: validation, schema-drift detection, reconciliation, alerting on the failures that matter rather than all of them.
  • Ship and run your own work — Docker, CI/CD, cloud environments. You deploy what you build.
  • Work directly with the client's engineering and AI teams on what the models actually need from the data — no layers in between.
  • Use AI-assisted development tools (Claude Code, Cursor, GitHub Copilot) as part of your daily workflow.

## What We're Looking For

  • 5+ years building and running production data pipelines
  • Strong Python and strong SQL — both at a level you'd defend in review
  • Real depth in third-party API integration: OAuth and token refresh, pagination strategies, rate limiting, retries and backoff, incremental syncs, and handling APIs whose documentation is wrong
  • Experience with a workflow orchestrator (Airflow, Dagster, Prefect, Temporal, or similar) and an opinion about it
  • Solid Postgres: schema design, query performance, migrations. Warehouse experience (BigQuery, Snowflake, Redshift) a plus
  • Batch and streaming ingestion patterns, and a sense of when each is the right call
  • Comfortable owning your own deployments — Docker, CI/CD, cloud (AWS or GCP)
  • A track record of shipping without a detailed spec and without process handed to you
  • Comfort working in the US timezone (overlapping working hours required)
  • Upper-Intermediate or higher English
  • Ability to own and drive features independently with minimal supervision

## Nice to have

  • Hands-on experience feeding LLM or RAG systems — chunking strategies, embedding pipelines, vector stores (pgvector, Qdrant, Pinecone, whatever you've used). Side projects count; commercial experience is not required here.
  • dbt, or another transformation layer you've run in anger
  • Data quality and observability tooling (Great Expectations, Monte Carlo, or your own)
  • Experience with managed connectors (Fivetran, Airbyte) — including knowing when to stop fighting them and write your own
  • Early-stage or agency background: greenfield work, shifting requirements, direct client contact

## What We Offer

  • Competitive salary above market average
  • Fully remote work
  • Flat structure — work directly with the client's engineering team, no middle management
  • Real technical ownership: you make the architecture decisions on the pipelines you build
  • Diverse and technically challenging projects
  • Modern AI-powered development workflow
  • Small team culture — your voice matters, your code ships
  • Paid vacation and sick leave
  • Flexible schedule within the US timezone overlap
  • Professional and career growth through real ownership, not courses and certificates

## How to Apply
Send your CV with a brief note about the nastiest API you've had to integrate — what made it hard, and how you made it reliable.

Skills

What Senior Data Engineering jobs ask for — and how much of it you have →
Apply

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available