Member of Technical Staff - Inference
Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at large scale to do the most challenging work.
In this role, you'll own token processing down to the lowest layers of the stack. You'll do things like: develop a new request scheduling strategy, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.
What you’ll do
Modify and extend state-of-the-art inference engines like vLLM and SGLang, and work on our own internal engine.
Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an nsys profile.
Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
Write and debug GPU kernels to excel in specific regimes, such as cascade attention
What we’re looking for
Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
Interest in MLSys research - great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.
Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
Great interpersonal and technical communication. Please don't use LLMs to write prose. We desk-reject slopful cover letters and resumes.
Interview process
Meet the CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. This is the first step because we respect your time.
Share an online whiteboard with a team member and work through a technical problem. We spend a lot of time at whiteboards, building intuition about complex systems together. It's a great way for us to see how you communicate technically, and a even better way for you to see what working at Sail is like.
Come in to Sail's SF office for an interview day. Meet the whole team, and work on a bunch of problems that closely simulates the work we do daily. We'll also ask you to give us a 20-30min 'chalk talk' about an interesting problem you've worked on before.
Offer. Once the team decides we want to work with you, we make a strong offer quickly and will be quite persistent over email/text/calls :)
Life at Sail
We work out of a beautiful, sunny office in downtown San Francisco. All meals are on us (and actually great; SF is a food paradise!). Everyone gets a Studio Display (or two) at their desk. We are serious about investing in anything that saves us time or energy. There are six different ways to make coffee or tea in the office. A friendly (hypoallergenic) black cat named Coco visits occasionally.
Skills
As published by ashby
Name, Email, Resume