Tech jobs
Job listings
2 jobs
23 hours ago
Research Engineer / Scientist - Storage for LLM
RemoteNorth America
United States
Research Engineer/Scientist focused on building distributed KV cache systems and GPU-aware caching layers for LLM inference. Core tasks include optimizing low-latency access, implementing memory-aware sharding, and integrating cache with token streaming pipelines.
batching c++ cache aware scheduling caching cuda +36 skills
23 hours ago
Research Engineer / Scientist - Storage for LLM
RemoteNorth America
United States
Research Engineer/Scientist to design and optimize distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on consistency, low-latency access, and memory-efficient sharding.
batching c++ cache aware scheduling caching algorithms cuda +27 skills