Senior Research Engineer / Scientist - Storage for LLM
Summary
Develops and optimizes distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on GPU-aware caching, consistency protocols, and performance tuning.
- Design distributed KV cache system
- Develop cache consistency and synchronization protocols
- Evaluate open source KV stores and build GPU aware caching layers
- Implement memory aware sharding eviction and replication
- Integrate cache with token streaming pipelines
- Monitor system performance and iterate caching algorithms
- Optimize caching latency throughput eviction
Perks/Benefits:
- 401k savings plan with company match
- Conference attendance
- Life insurance
- Long-term disability
- Medical/Dental/Vision insurance
- Open source contributions
- Paid Holidays
- Paid parental leave
- Paid personal time
- Paid sick days
- Research resources
- Short-term disability
- Wellbeing benefits