freehire launches on Product Hunt on 26 August.

Follow →

Research Engineer / Scientist - Storage for LLM

This position is no longer accepting applications(closed Aug 17, 2026).

Summary

Design and build distributed caching systems for large language models, optimizing GPU-aware KV caches, replication, and low-latency access to improve LLM performance.

- Build custom GPU aware caching layers - Design distributed KV cache system - Evaluate caching algorithms - Extend open source KV stores - Implement cache consistency protocols - Implement cache synchronization protocols - Implement eviction strategies - Implement memory aware sharding - Implement replication strategies - Integrate cache with batched decoding - Integrate cache with model parallelism - Integrate cache with token streaming pipelines - Monitor system performance - Optimize low-latency access and eviction policies - Store and retrieve intermediate states for transformer LLMs Perks/Benefits: - Conference attendance - Innovation-driven culture - Open source contributions - Research resources

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available