freehire launches on Product Hunt on 26 August.

Follow →

AI Infrastructure Engineer

Summary

Optimize LLM inference performance on Intel GPUs by profiling bottlenecks, writing custom kernels, and contributing to open-source serving frameworks like vLLM and SGLang.

We are looking for a performance-obsessed AI Infrastructure Engineer to push LLM inference to its absolute limits on Intel's next-generation GPU architectures. In this role, you will dive deep into the inference stack and redefine peak performance. You will work end-to-end across the stack: profiling bottlenecks, writing custom GPU kernels, and upstreaming your optimizations directly into industry-standard serving frameworks like vLLM and SGLang. Your optimizations will be instrumental in unloc…

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available