Senior Software Engineer - Distributed Data Platform
Summary
Build and run a distributed data platform handling ingestion, processing, and governance for ML/AI workloads using JVM, Kafka, and Kubernetes.
About The Role
We build and run the distributed platform behind our data ingestion, processing, and governance layer. It is a set of independently deployable services handling different types of data where correctness, lineage, and reliability matter as much as raw throughput.
We're hiring a Senior Engineer to own a meaningful slice of that platform. This is a hands-on software engineering role, not a pipeline (tool-configuration) role - you will design services, run them in production, and shape where the architecture goes next.
What You'll Do
- Own services in our data platform end to end - design, build, deploy, and operate them
- Build and evolve streaming and batch pipelines from ingestion through to governed, queryable data
- Make the platform debuggable at scale: SLOs, lineage, schema evolution, backpressure, replay, and failure recovery
- Contribute to architecture decisions - and argue against the ones you think are wrong
- Work alongside our ML and GenAI teams, whose workloads depend on the data this platform produces
What We're Looking For
- 5+ years building production backend or data systems, including time in a distributed or microservices environment
- Strong JVM (Java) engineering - most of our platform is JVM-based
- Solid SQL, plus hands-on experience with at least one columnar/analytical or NoSQL store
- Practical experience with distributed event streaming - Kafka or equivalent. Stream and batch processing frameworks (Spark, Flink etc) are a plus
- Comfort operating what you build: containers, Kubernetes, CI/CD, production debugging
- A structured, evidence-driven approach to problems - you reason about system behaviour rather than guessing at it
- BSc/MSc in Computer Science or equivalent practical experience
Preferred Qualifications
- Schema management and data contracts (Avro/Protobuf, schema registry, compatibility strategy)
- Change data capture and event-driven integration patterns
- Open table formats and lakehouse storage (Iceberg, Delta, Hudi), Parquet, object storage
- Data governance tooling - catalog, lineage, and data quality (OpenLineage, DataHub, OpenMetadata, or similar)
- Real-time analytical stores (ClickHouse, StarRocks, Druid, Trino)
- OpenTelemetry-based observability; Terraform or GitOps workflows
- Experience with data feeding ML, LLM, or retrieval/RAG workloads
If you're strong on the core and curious about the rest, we'd like to hear from you.
How We Work
- Small teams with real ownership - you'll be on the design decisions, not handed a ticket queue
- Code review and release responsibility in each development iteration
- We use AI coding assistants day to day. We expect fluency with them, and equally the judgment to know when not to lean on them - you own what ships either way.