Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)
Summary
The Inference Systems Backend Engineer will design and develop large-scale LLM training and inference systems, focusing on optimizing GPU cluster performance, model quantization, and elastic scheduling. The role involves collaborating with algorithm teams to integrate heterogeneous hardware and ensure high-concurrency reliability for large model architectures.
- Apply model quantization
- Build distributed LLM inference systems
- Collaborate with algorithm teams to optimize algorithms and systems
- Design large model training and inference architectures
- Develop large model training and inference systems
- Ensure high concurrency reliability and scalability
- Implement elastic scheduling
- Implement subgraph matching
- Improve compute utilization in distributed GPU clusters
- Integrate heterogeneous hardware with ML frameworks
- Manage GPU oversubscription
- Optimize model computation performance
- Orchestrate tasks
- Perform compiler optimization for models
- Schedule large scale inference traffic
- Tune large GPU training clusters