freehire launches on Product Hunt on 26 August.

Follow →

Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)

Summary

The Inference Systems Backend Engineer will design and develop large-scale LLM training and inference systems, focusing on optimizing GPU cluster performance, model quantization, and elastic scheduling. The role involves collaborating with algorithm teams to integrate heterogeneous hardware and ensure high-concurrency reliability for large model architectures.

- Apply model quantization - Build distributed LLM inference systems - Collaborate with algorithm teams to optimize algorithms and systems - Design large model training and inference architectures - Develop large model training and inference systems - Ensure high concurrency reliability and scalability - Implement elastic scheduling - Implement subgraph matching - Improve compute utilization in distributed GPU clusters - Integrate heterogeneous hardware with ML frameworks - Manage GPU oversubscription - Optimize model computation performance - Orchestrate tasks - Perform compiler optimization for models - Schedule large scale inference traffic - Tune large GPU training clusters

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available