AI Inference Engineer
Posted Updated
- Building and optimizing inference pipelines for large-scale model serving
- Working with frameworks like PyTorch, TensorRT, and vLLM to deploy models efficiently
- Implementing and optimizing ML models using techniques such as quantization, kernel fusion, and efficient batching
- Optimizing and implementing core ML operators such as GEMMs, convolutions, activations
- Investigating and resolving issues through system-level debugging and performance analysis
- Defining and applying practices for testing, deployment, and scaling AI systems
- BSc/MSc in Computer Science, Engineering, Mathematics, or related discipline
- Strong programming skills in C/C++ or Python in Linux environments using common development tools
- Solid knowledge of computer architecture, system software, data structures
- Hands-on experience implementing algorithms in high-level languages (C/C++/Python)
- Exposure to specialized hardware such as GPUs, FPGAs, DSPs, AI accelerators and frameworks such as OpenCL or CUDA
- Experience designing or working with high-performance software systems
- Solid knowledge of ML fundamentals
- Experience in model serving frameworks such as Triton Inference Server, DeepSpeed Inference, vLLM
- Experience with ML runtimes such as ONNX Runtime, TVM, IREE, XLA
- Experience deploying ML workloads across distributed systems
- Experience implementing and optimizing ML operators and kernels with focus on vectorization and efficient execution
- Experience in hardware-aware optimizations and performance tuning
- 2+ years of experience developing software targeting AI hardware
- Contribution to open-source projects such as LLVM/MLIR, PyTorch, TensorFlow, ONNX Runtime, xDSL, IREE