Member of Technical Staff - Inference
Summary
Build and optimize a multi-tenant LLM serving platform across cloud GPU fleets, including scheduling, failover, and integration with RL systems.
You will build infrastructure for efficient multi-tenant LLM serving across cloud GPU fleets, design GPU-aware scheduling and failover, optimize inference frameworks and parallelism, integrate distributed inference into RL systems, and establish CI/CD, observability, documentation, and incident response practices.
Responsibilities
- Build a multi-tenant LLM serving platform across cloud GPU fleets
- Design GPU-aware placement and scheduling algorithms
- Implement multi-region and multi-zone failover and traffic shifting
- Build autoscaling, routing, and load balancing
- Optimize model distribution and cold-start times
- Integrate and contribute to LLM inference frameworks
- Optimize parallelism, caching, memory management, quantization, and speculative decoding
- Profile kernels, memory bandwidth, and transport
- Develop reproducible performance suites
- Embed distributed inference into the RL stack
- Establish CI/CD with artifact promotion and performance gates
- Build observability and manage SLOs
- Document architectures and playbooks
- Mentor and collaborate cross-functionally
Requirements
- 3+ years building and operating large-scale ML or LLM services
- Experience with vLLM, SGLang, or TensorRT-LLM
- Familiarity with distributed and disaggregated serving infrastructure
- Understanding of prefill, decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism
- Experience debugging CUDA, NCCL, drivers, kernels, containers, service mesh, networking, and storage
- Python
- PyTorch
- AWS or GCP
- Kubernetes
- CUDA
- NCCL
- InfiniBand
Benefits
- Significant equity incentives
- Flexible work arrangement
- Full visa sponsorship
- Relocation support
- Professional development budget
- Regular team off-sites
- Conference attendance