freehire launches on Product Hunt on 26 August.

Follow →

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

Summary

Develops and optimizes high-performance inference engines for large AI models, focusing on GPU/NPU hardware acceleration, parallelism techniques, and performance tuning to reduce latency and improve throughput.

- Adapt inference engine for GPU and NPU hardware - Benchmark against vLLM and TensorRT-LLM - Develop distributed parallel inference solutions - Implement and iterate global large model inference solutions - Implement asynchronous stream scheduling - Implement mixture of experts expert parallelism - Implement pipeline parallelism - Implement sequence parallelism - Implement tensor parallelism - Improve GPU performance through operator fusion and compilation optimization - Optimize GPU memory access and compute pipeline - Optimize inference load balancing and parallel efficiency - Optimize large model inference engine architecture - Reduce cross card communication overhead

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available