Graduate Backend Inference Engineer — GPU & Optimization
Summary
Optimizes GPU/TPU-based inference engines for large models, focusing on performance, memory, and low-latency pipelines using C/C++, Python, and CUDA.
United States Digital Space LLC is seeking talented graduates or early-career engineers to join the Backend team in Singapore. The role focuses on iterating and optimizing the large model inference engine, with emphasis on GPU/TPU performance, memory management, and low-latency pipelines.
Ideal candidates have strong C/C++ and Python skills, CUDA experience, and a solid understanding of GPU architectures. The position involves cross-team collaboration and exposure to cutting-edge inference
