freehire launches on Product Hunt on 26 August. Follow the page and you'll hear the moment it opens.
The Inference Systems Backend Engineer will design and develop large-scale LLM training and inference systems, focusing on optimizing GPU cluster performance, model quantization, and elastic scheduling. The role involves collaborating with algorithm teams to integrate heterogeneous hardware and ensure high-concurrency reliability for large model architectures.
We couldn't check your fit for this role — add a CV to your profile to see it next time.