freehire launches on Product Hunt on 26 August.

Follow →

Graduate Backend Inference Engineer — GPU & Optimization

Summary

Optimizes GPU/TPU-based inference engines for large models, focusing on performance, memory, and low-latency pipelines using C/C++, Python, and CUDA.

United States Digital Space LLC is seeking talented graduates or early-career engineers to join the Backend team in Singapore. The role focuses on iterating and optimizing the large model inference engine, with emphasis on GPU/TPU performance, memory management, and low-latency pipelines.

Ideal candidates have strong C/C++ and Python skills, CUDA experience, and a solid understanding of GPU architectures. The position involves cross-team collaboration and exposure to cutting-edge inference

See also