Research Engineer / Scientist - Storage for LLM
Summary
Research Engineer/Scientist to design and optimize distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on consistency, low-latency access, and memory-efficient sharding.
- Design and implement distributed KV cache system
- Develop cache consistency and synchronization protocols
- Evaluate and extend open source KV stores
- Implement memory aware sharding eviction and replication
- Integrate cache with token streaming pipelines
- Monitor system performance and iterate caching algorithms
- Optimize low-latency access and eviction policies
Perks/Benefits:
- 401k match
- Conference attendance
- Life insurance
- Long-term disability
- Medical/Dental/Vision insurance
- Open source contributions
- Paid Holidays
- Paid parental leave
- Paid personal time
- Paid sick days
- Publishing Research
- Short-term disability
- Wellbeing benefits