Tech jobs
Job listings
2 jobs
21 hours ago
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
APAC
Singapore
Develops and optimizes high-performance inference engines for large AI models, focusing on GPU/NPU hardware acceleration, parallelism techniques, and performance tuning to reduce latency and improve throughput.
activation functions c# c++ computation graph computation graph optimization +32 skills
23 hours ago
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
RemoteNorth America
United States
Works on adapting inference engines to GPU/NPU hardware, benchmarking with vLLM/TensorRT-LLM, and optimizing distributed parallel inference solutions—including cache, memory, and latency improvements.
access optimization c# c++ computational graph computational graph optimization +28 skills