Tech jobs
Job listings
Machine Learning System Scheduling Engineer Graduate (Applied Machine Learning) - 2027 Start
Designs and optimizes scheduling systems for ML workloads, balancing compute, storage, and network resources across clusters.
LLM System Backend Engineer Graduate (AML Ark) - 2027 Start
Develops reinforcement learning infrastructure, RL training APIs, and optimizes LLM inference performance/cost for a multi-tenant, serverless platform.
Research Engineer / Scientist - Storage for LLM
Research Engineer/Scientist focused on building distributed KV cache systems and GPU-aware caching layers for LLM inference. Core tasks include optimizing low-latency access, implementing memory-aware sharding, and integrating cache with token streaming pipelines.
Large Language Model Inference System Engineer Graduate (Applied Machine Learning) - 2027 Start
Build and optimize inference systems for large language models as part of a Managed AI Service (MaaS), focusing on distributed KV caching, GPU performance tuning, and multi-node heterogeneous inference.
AI/LLM Network Software Development Engineer - San Jose
Develops next-gen AI/ML network infrastructure, co-designs host network apps for scalability/reliability, and designs high-speed network tech for AI/LLM apps, including monitoring and protocol stacks.
Software Development Engineer-AI/LLM Network-Global Frontier Tech Research Program-2027 Start
Develops and optimizes AI/ML network infrastructure, focusing on scalability, reliability, and performance for large-scale AI applications, including high-speed network technologies and communication frameworks.
Site Reliability Engineer, Machine Learning Systems - Singapore
Designs and maintains monitoring, disaster recovery, and resource management tools for large-scale ML systems, ensuring stable training, inference, and offline task execution across global data centers and cloud environments.
Senior Software Engineer- Metadata Storage
Design and build distributed metadata storage and services for infrastructure teams, focusing on key-value storage, coordination, and fault tolerance.
Software Development Engineer-AI/LLM Network-Global Frontier Tech Research Program-2027 Start (PhD)
Develops high-speed AI-native network infrastructure and protocols for LLM training and inference, optimizing host-network-application co-design.
AI/LLM Network Software Development Engineer Graduate (High Speed Network) - 2026 Start (PhD)
Develops high-speed network protocols and AI/ML communication frameworks to optimize host network performance for AI and LLM applications, including infrastructure co-design and diagnostic tools.
Senior Research Engineer / Scientist - Storage for LLM
Develops and optimizes distributed KV cache systems for large language model (LLM) token streaming pipelines, focusing on GPU-aware caching, consistency protocols, and performance tuning.