freehire launches on Product Hunt on 26 August.

Follow →

Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start

Summary

Builds and optimizes GPU/NPU inference engines for large AI models, focusing on latency, throughput, and parallelism techniques like tensor and pipeline parallelism.

- Adapt inference engine for GPU NPU hardware - Benchmark against TensorRT-LLM - Benchmark against vLLM - Design distributed parallel inference solutions - Implement Mixture of Experts parallelism - Implement compilation optimization - Implement operator fusion - Implement pipeline parallelism - Implement sequence parallelism - Implement tensor parallelism - Iterate large model inference engine architecture - Optimize GPU inference throughput - Optimize GPU memory access - Optimize cache performance - Optimize computing pipeline - Reduce inference latency - Schedule GPU streams asynchronously

See also