freehire launches on Product Hunt on 26 August.

Follow →

Senior AI Infra Engineer - Large Model Inference Systems (Multimodal / LLM / VLM)

Summary

Build and optimize inference infrastructure for large multimodal AI models, focusing on distributed serving, scheduling, and low-latency performance at massive scale.

Responsibilities About the Team We are dedicated to building the inference infrastructure for ultra-large-scale language models, vision-language models, and frontier multimodal AI systems. Our mission is to provide a robust, scalable, and high-performance foundation for distributed serving, heterogeneous scheduling, and low-latency inference at massive scale. You will work on some of the most challenging problems in large-model online serving, spanning traffic orchestration, throughput and late…

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available