Point your AI agent at freehire and let it find you a job.

Get the CLI →

Raydar

NewBe an early applicant

Member of Technical Staff, Inference Systems

Posted Updated 1 view
Discussion

Summary

Build the core LLM inference runtime for an AI infrastructure startup: batching, scheduling, routing, KV-cache/prefix caching, and multi-GPU model serving in Rust, plus profiling and architecture decisions with the founding team. Full-time onsite in San Francisco or the South Bay Area; $230k-$350k base plus equity.

About the company

Our client is an AI infrastructure company building high-performance model-serving technology. A small engineering team is developing a new inference runtime with ownership across the serving stack.

The role

Raydar is recruiting for this opportunity through the Paraform network. The position is with our client. Build the systems that determine inference latency, throughput, and cost. This Member of Technical Staff role covers runtime design, GPU serving, caching, and performance optimization.

What you'll do

- Build an inference runtime in Rust, including batching, scheduling, routing, and serving.

- Design KV-cache management, prefix caching, and related efficiency improvements.

- Scale model serving across multiple GPUs and nodes.

- Profile and benchmark the full inference pipeline.

- Make core architecture decisions with the founding team.

Requirements

What we're looking for

- 2 to 10 years of relevant systems engineering experience.

- Hands-on experience building or optimizing LLM inference and serving systems.

- Deep understanding of attention, KV cache, batching, and scheduling.

- Experience with production engines such as vLLM, SGLang, or TensorRT-LLM.

- Strong Rust, C++, Go, or systems-level Python/PyTorch skills; ability to become productive in Rust within 2 to 3 weeks if needed.

- Experience optimizing performance-critical backend or distributed systems.

Bonus points

- Production Rust experience or inference open-source contributions.

- CUDA or Triton kernel development.

- Multi-GPU or multi-node serving, speculative decoding, or prefill/decode disaggregation.

Benefits

Compensation and benefits

- Base salary: USD 230,000 to 350,000 per year.

- Equity: Competitive.

Location and work model

- South Bay Area or San Francisco, California, United States.

- Full-time onsite role.

- Visa transfers and new visa sponsorships are supported.

Skills

What Staff Software Engineering jobs ask for — and how much of it you have →

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available