AI Backend Engineer: Low-Latency, Scalable Inference API
Summary
Build and maintain low-latency, scalable inference APIs that power AI interactions in Doist’s product, optimizing for performance and reliability.
Doist in Singapore is seeking a Backend Engineer, AI to own the inference and orchestration layer powering AI interactions across the product. You will build and operate production systems delivering fast, reliable APIs between models and users.
You will design inference pipelines, manage monitoring and incident response, and push optimizations for latency, throughput, and cost. Collaborate with ML and frontend teams to ship robust features.