AI Infrastructure Engineer
Summary
Optimize LLM inference performance on Intel GPUs by profiling bottlenecks, writing custom kernels, and contributing to open-source serving frameworks like vLLM and SGLang.
We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance. You will work end-to-end across the stack: profiling bottlenecks, writing custom GPU kernels, and upstreaming your optimizations directly into industry-standard serving frameworks like vLLM and SGLang. Your optimizations will be instrumental in unloc…