Point your AI agent at freehire and let it find you a job.

Get the CLI →

Recrucial

New

Senior Data Engineer

Posted Updated 3 views
Discussion

Summary

Senior data engineer to own a ~5 TB/day event data platform end to end: streaming ingestion, ClickHouse/PostgreSQL storage and aggregation, reliability under failure, monitoring, and cost efficiency. Stack includes PostgreSQL, ClickHouse, Go, Kafka, AWS, and Terraform at a fully remote infrastructure product company.

Senior Data Engineer — Pipelines & Architecture

Remote · Full-time · Product Company

Who are we?

We are an infrastructure technology company building products that power data-driven businesses at global scale. Our ecosystem includes a global proxy network, managed APIs for large-scale web data collection, AI-powered automation tools, and a cross-platform Monetization SDK used across desktop, mobile, and smart TV environments. Our technology operates at significant scale, supporting millions of users and processing large volumes of traffic and telemetry every day. We are a founder-led, product-focused company backed by experienced entrepreneurs and technology leaders with a track record of building and scaling successful venture-backed businesses.

What do we do?

We build infrastructure that allows companies to collect, process, understand, and act on large volumes of data.

At the center of this role is our data platform.

Every request across our proxy network, every piece of SDK telemetry, and every newly instrumented product generates events. Those events ultimately become the numbers the business relies on for customer usage, billing, reporting, operational visibility, and decision-making. The data pipeline is expected to handle approximately 5 TB of events per day, and that volume will continue to grow. We are now building the architecture that can process this data reliably, efficiently, and predictably at scale.

Why do we do it?

At this scale, data infrastructure is not simply an internal analytics function. The accuracy of the pipeline directly affects what customers see, what we invoice, how we understand the health of our network, and how teams across the company make decisions. Our goal is therefore straightforward but technically demanding: make every important event fast to process, correct under failure, trustworthy to consume, and cost-efficient to store and query. We want teams to trust the data without needing to ask the engineer who built the pipeline whether a number is correct.

How do we do it?

We approach data engineering as a production engineering discipline rather than a collection of ETL scripts. That means thinking carefully about architecture, storage, partitioning, aggregation, observability, replayability, late-arriving data, schema evolution, backfills, reconciliation, and infrastructure cost. Our environment includes technologies such as PostgreSQL, ClickHouse, Go, Kafka or comparable streaming platforms, AWS, Terraform, and modern workflow and observability tooling. We are also an AI-forward engineering organization. The team actively uses LLMs and agentic coding tools as part of everyday engineering work. Most importantly, engineers here are expected to own problems end to end: understand the business requirement, design the architecture, build it, operate it in production, and improve it when reality inevitably disagrees with the original design.

Who works for us?

You will join a small, senior, fully remote engineering team working closely across backend engineering, proxy infrastructure, and product. The organization is intentionally flat and low-process. We value autonomy, technical judgment, clear communication, and people who can move from identifying a problem to implementing a solution without layers of coordination. English is our working language, and much of our collaboration is asynchronous, so being able to communicate technical reasoning clearly is important.

Why is there a vacancy?

The scale and importance of our data infrastructure are growing. We need a senior engineer who can take ownership of the data pipeline as a system—not simply maintain individual jobs or databases. This person will help take the streaming architecture from design into production, establish the storage and aggregation strategy, improve reliability under failure, and ensure the platform remains operationally and economically sustainable as traffic increases. This is an opportunity to shape the architecture rather than inherit a finished system.

Who are we looking for?

  • We are looking for an experienced Senior Data Engineer who has already built and operated business-critical data infrastructure in production.
  • You should be comfortable owning architectural decisions as well as writing and debugging production code.
  • You understand that at billions of rows and multi-terabyte daily ingestion volumes, seemingly small decisions around partitioning, retention, indexing, aggregation, or replay behavior can have very large consequences.

What professional skills are important to us?

We are looking for someone with:

  • 7+ years of experience building and operating production data pipelines.

  • Strong hands-on PostgreSQL expertise, including partitioning, indexing, bulk-insert performance, and experience understanding when PostgreSQL is no longer the appropriate solution.

  • Production experience working with databases exceeding 5 billion records.

  • Strong hands-on ClickHouse experience, including engine selection, partitioning, materialized views, and query optimization across billions of rows.

  • Deep knowledge of SQL and analytical data modeling.

  • Experience designing trustworthy aggregation layers from raw event data.

  • Experience with high-volume telemetry, observability, or event data, ideally involving multi-terabyte daily ingestion.

  • Proven ownership of critical production data pipelines.

  • Experience planning and scaling systems while balancing technical requirements, infrastructure cost, and available engineering resources.

  • Strong written and verbal communication skills.

  • Ability to operate effectively in an autonomous, fast-moving startup environment.

  • Active experience using agentic coding and modern AI development tools.

Nice to have

It would be particularly valuable if you also have experience with:

  • Kafka or another production streaming platform, including partitioning, consumer groups, and delivery guarantees.

  • Go, as our pipeline services are written in Go.

  • Usage-based metering or billing pipelines.

  • Reconciliation and late-arriving data handling.

  • Workflow orchestration tools such as Temporal, Airflow, or Dagster.

  • Transformation, BI, and observability tooling such as dbt, Metabase, or Grafana.

  • AWS, particularly RDS, Lambda, SQS, S3/Parquet, and EC2.

  • Terraform.

  • Self-hosting and operating ClickHouse or Kafka clusters rather than exclusively consuming managed services.

What will you do in the project?

You will own major parts of the data platform from architecture through production operations.

Your responsibilities will include:

Taking the streaming pipeline from design to production.

Designing the storage architecture, including how events are structured, partitioned, retained, and expired.

Building ingestion that remains correct under failure, including replays, duplicates, out-of-order events, and late-arriving data.

Owning the aggregation layer used by customer-facing usage views and internal reporting.

Defining reliable data models, metrics, and freshness guarantees.

Maintaining predictable performance across both high-volume ingestion and analytical queries as traffic grows.

Automating schema changes, backfills, reconciliation, and other operational data workflows.

Replacing manual production interventions with repeatable, observable processes.

Building monitoring around freshness, pipeline lag, ingestion failures, data quality, and infrastructure cost.

Creating alerts that surface problems before incorrect data reaches customers or business dashboards.

Participating in on-call and incident response for the data platform and driving technical follow-ups after incidents.

Collaborating with backend, infrastructure, and product teams to ensure that the events we collect actually answer the questions the business needs to ask.

You will not simply be handed tickets for individual pipelines. We expect you to influence how the data platform itself should work.

What conditions do we offer?

  • Full-time, fully remote cooperation.

  • Flexible working hours with a preference for at least 4 hours of overlap with GMT+2.

  • 3-month trial period.

  • Unlimited vacation and sick days, with advance notice for planned leave.

  • Competitive compensation.

  • Equity incentive plan.

  • Health insurance reimbursement.

  • Remote-work setup compensation.

  • A small, experienced team with direct ownership and minimal bureaucracy.

  • The opportunity to make architectural decisions that have visible impact across the entire business.

What is our hiring process?

We keep the process focused on conversations, practical engineering judgment, and real-world problem solving rather than artificial interview exercises.

1. Intro Call — approximately 30 minutes
An initial conversation to discuss your background, expectations, and the opportunity.

2. Culture Interview — approximately 30–60 minutes
A conversation focused on ownership, collaboration, communication, and how you approach work in a remote startup environment.

3. Two Technical Interviews — approximately 60–90 minutes each
Deep technical discussions around architecture, data pipelines, production systems, scalability, reliability, and the engineering decisions you have made in real environments.

4. Reference Checks — 2–3 references

Skills

See also

Data Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available