Tech jobs
Job listings
1 job
yesterday
Senior Research Engineer / Scientist - Storage for LLM
RemoteNorth AmericaSenior
United States
Develops and optimizes distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on GPU-aware caching, consistency protocols, and performance tuning.
autoregressive decoding batched decoding c++ cache consistency cuda +34 skills