Research Engineer / Scientist - Storage for LLM
Summary
Research Engineer/Scientist focused on building distributed KV cache systems and GPU-aware caching layers for LLM inference. Core tasks include optimizing low-latency access, implementing memory-aware sharding, and integrating cache with token streaming pipelines.
- Build custom GPU aware caching layers
- Collaborate with inference and serving teams
- Design and implement distributed KV cache system
- Develop cache consistency and synchronization protocols
- Evaluate and extend open source KV stores
- Implement memory aware sharding eviction and replication
- Integrate cache with token streaming pipelines
- Monitor system performance and iterate on caching algorithms
- Optimize low-latency access and eviction policies
Perks/Benefits:
- 401k matching
- Life insurance
- Long-term disability
- Medical, dental, and vision insurance
- Paid Holidays
- Paid parental leave
- Paid personal time
- Paid sick days
- Short-term disability
- Wellbeing benefits