Tech jobs
Job listings
2 jobs
yesterday
Research Engineer / Scientist - Storage for LLM
RemoteNorth America
United States
Research Engineer/Scientist to design and optimize distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on consistency, low-latency access, and memory-efficient sharding.
batching c++ cache aware scheduling caching algorithms cuda +27 skills
yesterday
Senior Research Engineer / Scientist - Storage for LLM
RemoteNorth AmericaSenior
United States
Develops and optimizes distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on GPU-aware caching, consistency protocols, and performance tuning.
autoregressive decoding batched decoding c++ cache consistency cuda +34 skills