freehire launches on Product Hunt on 26 August.

Follow →

Lead Data Engineer

Summary

Build a streaming-first data platform for autonomous cleaning robots, designing medallion architecture and semantic layers to turn real-time telemetry into trusted fleet KPIs.

Lead Data Engineer

Job type: Full Time · Department: Research & Development (R&D) · Work type: On-Site

Singapore, Singapore

Your data sources have wheels. LionsBot designs and builds autonomous cleaning robots that work in the real world: malls, airports, offices and industrial sites across 30+ countries. More than 5,000 of them stream telemetry to us in real time: missions, maps, locations, incidents, battery health. We move fast: small team, quick decisions, zero bureaucracy, hardware you can kick.

You’ll be our first dedicated data hire, a true 0→1, greenfield ownership role. The foundations are in place: real-time telemetry streams from the fleet into a time-series store, with dashboards on top. Our fleet has now grown to the point where data deserves a full-time owner, so we’re making it a first-class function. End-to-end, it’s yours.

The mission: take us from "a pipeline that works" to a streaming-first data platform with a proper medallion architecture: bronze raw telemetry, silver cleaned and conformed, gold business-ready marts. On top of it all, a semantic layer where every metric has exactly one definition and everyone trusts the number.

The fun problems, all real

  • Robots report cumulative lifetime odometers on every mission row. Sum the wrong column and your fleet total inflates 1,000×. Design the models that make that mistake impossible.

  • A sensor glitch claims one robot cleaned 2.5 million m² in twenty minutes. Build the data quality and anomaly detection that catches it before a human ever sees it.

  • Robots in basements with bad Wi‑Fi send late‑arriving, out‑of‑order data. Make the pipelines idempotent anyway.

  • Real‑time fleet health: which robots are sick right now, across 30+ countries and time zones?

What you will do

  • Own the data platform end-to-end: ingestion, storage, modeling, serving, dashboards. Real-time event streams from the fleet land in a time-series database today. Where it goes next is your call.

  • Design the medallion architecture: bronze, silver and gold layers over high-volume IoT telemetry, with clear data contracts agreed with the backend teams so quality is designed in at the source.

  • Build the metrics/semantic layer: canonical, documented, version‑controlled definitions for fleet KPIs: cleaning hours, area, mission success, incident rates, robot health. A genuine single source of truth.

  • Run data quality & observability like production software: freshness SLAs, validation, dedup, outlier handling, anomaly alerts. Flag the weird number before leadership does.

  • Design, tune and re‑architect databases at scale: schemas, indexes, continuous aggregates, compression, downsampling, retention and partitioning, treating them like the production systems they are.

  • Make analytics self‑serve: dashboards and models for ops, product, leadership and OEM partners, plus fast, rigorous answers to the high‑stakes ad‑hoc questions.

  • Shape the roadmap: we run lean today, so what comes next is genuinely open: OLAP, orchestration, transformation tooling, lakehouse patterns. You evaluate, make the case, and we adopt what earns its keep.

  • Work AI‑native: we pair humans with LLM‑powered analytics agents daily. You’ll design the platform so both humans and AI agents can query it safely and correctly.

What we are looking for

  • 3+ years working with data in production: data engineering, analytics engineering, or backend with heavy data exposure. We hire for trajectory, not year count.

  • Strong SQL, solid PostgreSQL and confident database design: schemas, indexes and data models that hold up as data grows. Time‑series databases like TimescaleDB or InfluxDB are a big plus, but you’ll learn them fast here.

  • Comfortable with event‑driven data: you’ve worked with streaming or message‑queue systems like Kafka, or you’re a data‑minded backend engineer keen to go deeper on real‑time.

  • Solid Python for pipelines and tooling, and comfortable reading Go or Java services.

  • You’ve shipped dashboards and metrics people actually used, whatever the BI tool.

  • Fast and autonomous, like our robots: high ownership, pragmatic trade‑offs, comfortable with ambiguity, ships iteratively.

  • Clear communication: you translate data into decisions, not just charts.

Nice to have:

  • Production streaming chops: you know your at‑least‑once from your exactly‑once.

  • IoT, robotics, or high‑volume device telemetry experience.

  • AWS, especially EKS, RDS and S3, with exposure to Azure or GCP.

  • OLAP engines, orchestration or transformation tooling, CDC pipelines.

  • Search engines like Quickwit or Elasticsearch, or graph databases like Neo4j.

  • Geospatial data: our robots navigate real floors, so maps and location streams are first‑class citizens.

  • Experience making data platforms LLM/agent‑friendly: semantic layers, governed self‑serve.

  • Familiarity with OpenRMF, ROS or robotics‑related communication stacks

Why join:

  • 0→1 ownership. The architecture, the standards, and the tooling choices are yours to shape. Eventually, so is the team.

  • Grow with the function. We’re hiring for trajectory: as data grows from one person into a team, you’re first in line to lead it.

  • Physical‑world data at real scale. Thousands of robots, 30+ countries, real‑time streams. Not clickstream. Not ad attribution. Robots.

  • Visible impact. Small team, direct line to leadership. What you ship this week is used in decisions next week.

  • AI‑forward team. We already run AI‑assisted analytics in production workflows. You’ll multiply it, not fight it.

See also