freehire launches on Product Hunt on 26 August.

Follow →

Performance Engineer, Inference

- Build and train speculative decoders - Collaborate on architecture co design - Integrate model and kernel artifacts into serving stack - Measure and tune acceptance rate - Modify inference runtime schedulers and allocators - Operate distributed serving across nodes - Optimize KV cache internals - Own production inference serving path end to end - Produce and defend latency and throughput metrics - Profile performance and trace bottlenecks - Provide on call inference SLO ownership

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available