Member of Technical Staff, Inference Systems
Summary
Build the core LLM inference runtime for an AI infrastructure startup: batching, scheduling, routing, KV-cache/prefix caching, and multi-GPU model serving in Rust, plus profiling and architecture decisions with the founding team. Full-time onsite in San Francisco or the South Bay Area; $230k-$350k base plus equity.
About the company
Our client is an AI infrastructure company building high-performance model-serving technology. A small engineering team is developing a new inference runtime with ownership across the serving stack.
The role
Raydar is recruiting for this opportunity through the Paraform network. The position is with our client. Build the systems that determine inference latency, throughput, and cost. This Member of Technical Staff role covers runtime design, GPU serving, caching, and performance optimization.
What you'll do
- Build an inference runtime in Rust, including batching, scheduling, routing, and serving.
- Design KV-cache management, prefix caching, and related efficiency improvements.
- Scale model serving across multiple GPUs and nodes.
- Profile and benchmark the full inference pipeline.
- Make core architecture decisions with the founding team.
Requirements
What we're looking for
- 2 to 10 years of relevant systems engineering experience.
- Hands-on experience building or optimizing LLM inference and serving systems.
- Deep understanding of attention, KV cache, batching, and scheduling.
- Experience with production engines such as vLLM, SGLang, or TensorRT-LLM.
- Strong Rust, C++, Go, or systems-level Python/PyTorch skills; ability to become productive in Rust within 2 to 3 weeks if needed.
- Experience optimizing performance-critical backend or distributed systems.
Bonus points
- Production Rust experience or inference open-source contributions.
- CUDA or Triton kernel development.
- Multi-GPU or multi-node serving, speculative decoding, or prefill/decode disaggregation.
Benefits
Compensation and benefits
- Base salary: USD 230,000 to 350,000 per year.
- Equity: Competitive.
Location and work model
- South Bay Area or San Francisco, California, United States.
- Full-time onsite role.
- Visa transfers and new visa sponsorships are supported.
Skills
As published by workable · 3 questions
Basics
First name, Last name, Email, Headline, Phone, Address, Photo, Education, Experience, Summary, Resume, Cover letter
Short answers (1)
- Public LinkedIn profile URL (https://www.linkedin.com/in/...)
Pick from a list (2)
- Will you require work authorization of any kind?
- Are you able to work onsite in the South Bay Area or San Francisco?
