Performance Engineer, Inference
- Build and train speculative decoders
- Collaborate on architecture co design
- Integrate model and kernel artifacts into serving stack
- Measure and tune acceptance rate
- Modify inference runtime schedulers and allocators
- Operate distributed serving across nodes
- Optimize KV cache internals
- Own production inference serving path end to end
- Produce and defend latency and throughput metrics
- Profile performance and trace bottlenecks
- Provide on call inference SLO ownership