freehire launches on Product Hunt on 26 August.

Follow →

Senior Inference Runtime Engineer

You will own the performance-critical serving layer for self-hosted large language models. You will optimize scheduling, batching, KV cache behavior, decoding, and streaming; tune inference runtimes; profile GPU, network, tokenizer, proxy, and worker bottlenecks; lead model onboarding; define runtime playbooks and safe defaults; and turn benchmark findings into production improvements.

Responsibilities

  • Optimize prefill and decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming
  • Tune and operate LLM inference runtimes for latency, throughput, GPU utilization, and cost efficiency
  • Profile bottlenecks across GPU memory, HBM bandwidth, NCCL, networking, tokenizers, proxies, and model workers
  • Lead model onboarding and select runtime, parallelism, quantization, context length, and rollback strategies
  • Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal workloads, prompt caching, and provider parameters
  • Partner with SRE and performance and evaluation engineers on production runtime improvements

Requirements

  • 6+ years in systems, ML infrastructure, or high-performance backend engineering
  • Experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton
  • Strong understanding of GPU memory, CUDA, NCCL, KV cache, batching, streaming, and distributed inference
  • Go or Python proficiency
  • Ability to read runtime source code and profiling traces
  • Experience operating production inference services with latency, availability, and cost targets
  • Ability to translate performance work into reliability, latency, and margin improvements

Benefits

  • Welfare benefits

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available