Tech jobs
Job listings
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Develop and optimize a high-performance GPU inference engine for large language models, focusing on operator fusion, compilation, and distributed parallelism to reduce latency and boost throughput.
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Build and optimize ByteDance’s large-model inference engine, focusing on GPU performance, operator fusion, compilation optimizations, and distributed parallelism to reduce latency and boost throughput.
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Optimize and architect ByteDance’s large-model inference engine for GPU/NPU, using CUDA, operator fusion, and distributed parallelism to boost throughput and cut latency.
Software Engineer, AI Infrastructure & Operations, AI Practice
Build and maintain secure, scalable AI infrastructure for Singapore’s government, automating ML pipelines, optimizing LLMs, and enforcing compliance standards.
Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)
Build and optimize large-scale distributed inference systems for LLMs on VolcanoEngine, tuning GPU clusters and scheduling traffic to handle hundreds of billions of tokens daily.
Senior Backend Engineer - Machine Learning Platform (R&D, CTR/VTR Predictor) - Ego team
Build and optimize low-latency, high-throughput ML inference services for CTR/CVR prediction and generative recommendation using LLMs, focusing on GPU acceleration and end-to-end pipeline optimization.
Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)
Build and optimize large-scale distributed inference systems for LLMs on VolcanoEngine, tuning GPU clusters and traffic scheduling to handle hundreds of billions of tokens daily.
Backend Engineer - AML Framework Development (Search, Ads, and Recommendation Direction)
Build and optimize a high-performance GPU inference engine for large AI models powering ads, search, and recommendation ranking systems using C/C++, Python, and CUDA.
Backend Engineer (Recommendation System/ LLM)
Build and scale backend infrastructure for AI-driven recommendation systems using ML serving platforms, Kubernetes, and distributed systems to deliver personalized financial services.
Backend Engineer (Recommendation System/ LLM)
Builds and scales backend infrastructure for AI-driven recommendation systems and LLM serving platforms in a fintech company.
Senior Computer Vision Engineer
Build and optimize real-time computer vision models (YOLO, tracking) in PyTorch/TensorFlow to analyze video streams for retail and operational insights.
Machine Learning Engineer
Build and optimize large-scale ML training and inference pipelines for low-latency trading systems using PyTorch, CUDA, and GPU acceleration.
Machine Learning Engineer
Build and optimize distributed ML training/inference pipelines for low-latency trading systems using PyTorch, CUDA, and GPU acceleration.
Senior AI/LLM Engineer
Lead training, alignment, and optimization of large language models using RLHF, SFT, and quantization; build reward models, red-team models, and optimize inference pipelines in Python/C++/Rust.
Machine Learning Engineer
Build and deploy computer-vision and generative-AI models to track construction progress from site imagery, integrating ML into Timescapes’ SaaS for global contractors.
Machine Learning Lead (Computer Vision)
Lead the design and deployment of AI/ML solutions using Azure ML, OpenAI APIs, and Generative AI for computer vision tasks like object detection and retail shelf analysis.
Machine Learning Engineer
Build and deploy computer-vision and generative-AI models to track construction progress, validate claims, and improve site visibility for large firms.
Senior Machine Learning Engineer
Senior ML engineer building perception systems for autonomous maritime robots using computer vision, sensor fusion, and real-time ML pipelines.
Machine Learning Engineer
Senior DevOps ML Engineer builds and runs AI platforms, splitting time between GPU-accelerated Kubernetes, backend services, and MLOps to productionize LLMs and digital avatars.
Staff Python / PyTorch Developer — Frontend Inference Compiler – Dubai
Build and optimize high-performance inference APIs and ML features for generative AI models running on custom hardware using Python, PyTorch, and C++.