freehire launches on Product Hunt on 26 August.

Follow →

Senior Software Engineer - Distributed Data Platform

About The Role

We build and run the distributed platform behind our data ingestion, processing, and governance layer. It is a set of independently deployable services handling different types of data where correctness, lineage, and reliability matter as much as raw throughput.

We're hiring a Senior Engineer to own a meaningful slice of that platform. This is a hands-on software engineering role, not a pipeline (tool-configuration) role - you will design services, run them in production, and shape where the architecture goes next.

What You'll Do

  • Own services in our data platform end to end - design, build, deploy, and operate them
  • Build and evolve streaming and batch pipelines from ingestion through to governed, queryable data
  • Make the platform debuggable at scale: SLOs, lineage, schema evolution, backpressure, replay, and failure recovery
  • Contribute to architecture decisions - and argue against the ones you think are wrong
  • Work alongside our ML and GenAI teams, whose workloads depend on the data this platform produces

What We're Looking For

  • 5+ years building production backend or data systems, including time in a distributed or microservices environment
  • Strong JVM (Java) engineering - most of our platform is JVM-based
  • Solid SQL, plus hands-on experience with at least one columnar/analytical or NoSQL store
  • Practical experience with distributed event streaming - Kafka or equivalent. Stream and batch processing frameworks (Spark, Flink etc) are a plus
  • Comfort operating what you build: containers, Kubernetes, CI/CD, production debugging
  • A structured, evidence-driven approach to problems - you reason about system behaviour rather than guessing at it
  • BSc/MSc in Computer Science or equivalent practical experience

Preferred Qualifications

  • Schema management and data contracts (Avro/Protobuf, schema registry, compatibility strategy)
  • Change data capture and event-driven integration patterns
  • Open table formats and lakehouse storage (Iceberg, Delta, Hudi), Parquet, object storage
  • Data governance tooling - catalog, lineage, and data quality (OpenLineage, DataHub, OpenMetadata, or similar)
  • Real-time analytical stores (ClickHouse, StarRocks, Druid, Trino)
  • OpenTelemetry-based observability; Terraform or GitOps workflows
  • Experience with data feeding ML, LLM, or retrieval/RAG workloads

If you're strong on the core and curious about the rest, we'd like to hear from you.

How We Work

  • Small teams with real ownership - you'll be on the design decisions, not handed a ticket queue
  • Code review and release responsibility in each development iteration
  • We use AI coding assistants day to day. We expect fluency with them, and equally the judgment to know when not to lean on them - you own what ships either way.

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available