Tech jobs
Job listings
Research Engineer / Scientist - Storage for LLM
Research Engineer/Scientist focused on building distributed KV cache systems and GPU-aware caching layers for LLM inference. Core tasks include optimizing low-latency access, implementing memory-aware sharding, and integrating cache with token streaming pipelines.
AI/LLM Network Software Development Engineer - San Jose
Develops next-gen AI/ML network infrastructure, co-designs host network apps for scalability/reliability, and designs high-speed network tech for AI/LLM apps, including monitoring and protocol stacks.
Research Engineer - LLM Training Infrastructure - Seed Infra
Research Engineer focused on optimizing and scaling infrastructure for large language model (LLM) training, addressing performance bottlenecks and designing distributed training strategies for exascale systems.
Software Development Engineer-AI/LLM Network-Global Frontier Tech Research Program-2027 Start
Develops and optimizes AI/ML network infrastructure, focusing on scalability, reliability, and performance for large-scale AI applications, including high-speed network technologies and communication frameworks.
Software Development Engineer-AI/LLM Network-Global Frontier Tech Research Program-2027 Start (PhD)
Develops high-speed AI-native network infrastructure and protocols for LLM training and inference, optimizing host-network-application co-design.
AI/LLM Network Software Development Engineer Graduate (High Speed Network) - 2026 Start (PhD)
Develops high-speed network protocols and AI/ML communication frameworks to optimize host network performance for AI and LLM applications, including infrastructure co-design and diagnostic tools.
Research Engineer / Scientist - Storage for LLM
Research Engineer/Scientist to design and optimize distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on consistency, low-latency access, and memory-efficient sharding.
Senior Research Engineer / Scientist - Storage for LLM
Develops and optimizes distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on GPU-aware caching, consistency protocols, and performance tuning.