Tech jobs
Job listings
2 jobs
yesterday
Senior Research Engineer / Scientist - Storage for LLM
RemoteNorth AmericaSenior
United States
Develops and optimizes distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on GPU-aware caching, consistency protocols, and performance tuning.
autoregressive decoding batched decoding c++ cache consistency cuda +34 skills
yesterday
Research Engineer - LLM/VLM Inference Optimization (Seed Infra)
RemoteNorth America
United States
Design and build high-performance inference engines for large language and vision-language models, optimizing latency, throughput, and cost while collaborating with research teams.
c plus plus c# containerization conv2d cpu gpu +25 skills