Senior Inference Runtime Engineer
You will own the performance-critical serving layer for self-hosted large language models. You will optimize scheduling, batching, KV cache behavior, decoding, and streaming; tune inference runtimes; profile GPU, network, tokenizer, proxy, and worker bottlenecks; lead model onboarding; define runtime playbooks and safe defaults; and turn benchmark findings into production improvements.
Responsibilities
- Optimize prefill and decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming
- Tune and operate LLM inference runtimes for latency, throughput, GPU utilization, and cost efficiency
- Profile bottlenecks across GPU memory, HBM bandwidth, NCCL, networking, tokenizers, proxies, and model workers
- Lead model onboarding and select runtime, parallelism, quantization, context length, and rollback strategies
- Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal workloads, prompt caching, and provider parameters
- Partner with SRE and performance and evaluation engineers on production runtime improvements
Requirements
- 6+ years in systems, ML infrastructure, or high-performance backend engineering
- Experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton
- Strong understanding of GPU memory, CUDA, NCCL, KV cache, batching, streaming, and distributed inference
- Go or Python proficiency
- Ability to read runtime source code and profiling traces
- Experience operating production inference services with latency, availability, and cost targets
- Ability to translate performance work into reliability, latency, and margin improvements
Benefits
- Welfare benefits