Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Summary
Builds and optimizes GPU/NPU inference engines for large AI models, focusing on latency, throughput, and parallelism techniques like tensor and pipeline parallelism.
freehire launches on Product Hunt on 26 August. Follow the page and you'll hear the moment it opens.
Follow →Builds and optimizes GPU/NPU inference engines for large AI models, focusing on latency, throughput, and parallelism techniques like tensor and pipeline parallelism.