Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Summary
Develops and optimizes high-performance inference engines for large AI models, focusing on GPU/NPU hardware acceleration, parallelism techniques, and performance tuning to reduce latency and improve throughput.
- Adapt inference engine for GPU and NPU hardware
- Benchmark against vLLM and TensorRT-LLM
- Develop distributed parallel inference solutions
- Implement and iterate global large model inference solutions
- Implement asynchronous stream scheduling
- Implement mixture of experts expert parallelism
- Implement pipeline parallelism
- Implement sequence parallelism
- Implement tensor parallelism
- Improve GPU performance through operator fusion and compilation optimization
- Optimize GPU memory access and compute pipeline
- Optimize inference load balancing and parallel efficiency
- Optimize large model inference engine architecture
- Reduce cross card communication overhead