freehire launches on Product Hunt on 26 August.

Follow →

Research Engineer / Scientist - Storage for LLM

Summary

Research Engineer/Scientist focused on building distributed KV cache systems and GPU-aware caching layers for LLM inference. Core tasks include optimizing low-latency access, implementing memory-aware sharding, and integrating cache with token streaming pipelines.

- Build custom GPU aware caching layers - Collaborate with inference and serving teams - Design and implement distributed KV cache system - Develop cache consistency and synchronization protocols - Evaluate and extend open source KV stores - Implement memory aware sharding eviction and replication - Integrate cache with token streaming pipelines - Monitor system performance and iterate on caching algorithms - Optimize low-latency access and eviction policies Perks/Benefits: - 401k matching - Life insurance - Long-term disability - Medical, dental, and vision insurance - Paid Holidays - Paid parental leave - Paid personal time - Paid sick days - Short-term disability - Wellbeing benefits

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available