Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Summary
Works on adapting inference engines to GPU/NPU hardware, benchmarking with vLLM/TensorRT-LLM, and optimizing distributed parallel inference solutions—including cache, memory, and latency improvements.
- Adapt inference engine to GPU and NPU hardware
- Benchmark with vLLM and TensorRT-LLM
- Design distributed parallel inference solutions
- Develop cache optimization strategies
- Handle load imbalance and parallel efficiency
- Implement inference framework performance optimizations
- Implement mixture of experts expert parallelism
- Implement operator fusion
- Implement pipeline parallelism
- Implement sequence parallelism
- Implement stream asynchronous scheduling
- Implement tensor parallelism
- Improve end to end GPU performance
- Increase single card inference throughput
- Optimize GPU memory access
- Optimize computing pipeline
- Optimize large model inference engine architecture
- Optimize multi card splitting and deployment
- Perform compilation optimization
- Reduce cross card communication overhead
- Reduce inference latency
Perks/Benefits:
- 401k match
- Dental insurance
- Life insurance
- Long-term disability
- Medical insurance
- Paid Holidays
- Paid parental leave
- Paid personal time
- Paid sick days
- Short-term disability
- Vision insurance
- Wellbeing benefits