Tech jobs
Job listings
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Develop and optimize a high-performance GPU inference engine for large language models, focusing on operator fusion, compilation, and distributed parallelism to reduce latency and boost throughput.
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Build and optimize ByteDance’s large-model inference engine, focusing on GPU performance, operator fusion, compilation optimizations, and distributed parallelism to reduce latency and boost throughput.
Backend Software Engineer - Recommendation Content Understanding Architecture
Build and optimize TikTok’s multi-modal content understanding pipeline, leveraging vector retrieval and large-scale ML models to improve recommendation relevance and system performance.
Backend Inference Runtime Engineer Graduate (AML Inference) - 2027 Start
Optimize and architect ByteDance’s large-model inference engine for GPU/NPU, using CUDA, operator fusion, and distributed parallelism to boost throughput and cut latency.
Graduate Backend Inference Engineer — GPU & Optimization
Optimizes GPU/TPU-based inference engines for large models, focusing on performance, memory, and low-latency pipelines using C/C++, Python, and CUDA.
Backend Engineer - Model Training Infra (Singapore) Technology - Backend Singapore Regular
Build and scale the ML infrastructure powering ByteDance’s global ad, search, and e-commerce ranking systems using C/C++, CUDA, Python, and frameworks like TensorFlow/PyTorch.
Backend Engineer - Multi-Modal Content & Recommendations
Design and improve TikTok’s multi-modal content understanding pipeline for recommendations, ensuring high availability and scalable vector processing.
Backend Engineer Intern (ByteRec Recommendation Infrastructure) - 2026 Start (BS/MS)
Backend engineering intern building scalable recommendation infrastructure using multi-modal content processing, vector retrieval, and RAG systems in Python/C++.
Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)
Build and optimize large-scale distributed inference systems for LLMs on VolcanoEngine, tuning GPU clusters and scheduling traffic to handle hundreds of billions of tokens daily.
Senior Backend Engineer - Machine Learning Platform (R&D, CTR/VTR Predictor) - Ego team
Build and optimize low-latency, high-throughput ML inference services for CTR/CVR prediction and generative recommendation using LLMs, focusing on GPU acceleration and end-to-end pipeline optimization.
Backend Engineer - Model Training Infra (Singapore)
Build and scale the ML infrastructure that trains and serves ranking models for ads, search, and feeds using C/C++/CUDA/Python and frameworks like TensorFlow/PyTorch.
LLM System Backend Engineer Graduate (AML Ark) - 2027 Start
Build and optimize distributed LLM training and inference platforms for ByteDance’s Volcano Ark, focusing on GPU clusters, RL workflows, and scalable agent infrastructure.
LLM System Backend Engineer Graduate (AML Ark) - 2027 Start
Build and optimize distributed training and inference platforms for large language models, including serverless post-training, reinforcement learning infrastructure, and high-performance GPU inference systems.
Backend Engineer - Model Training Infra (Singapore)
Build and scale AI infrastructure for ranking models in ads, search, and e-commerce using C/C++/CUDA/Python and deep learning frameworks like TensorFlow/PyTorch.
Inference Systems Backend Engineer - ARK Large Model Platform (Singapore)
Build and optimize large-scale distributed inference systems for LLMs on VolcanoEngine, tuning GPU clusters and traffic scheduling to handle hundreds of billions of tokens daily.
Backend Engineer - AML Framework Development (Search, Ads, and Recommendation Direction)
Build and optimize a high-performance GPU inference engine for large AI models powering ads, search, and recommendation ranking systems using C/C++, Python, and CUDA.
Senior Devops Engineer
Build and automate CI/CD pipelines for deploying LLMs and AI agents on Kubernetes in a regulated banking environment using Docker, Terraform, and model serving platforms like vLLM.
DevOps Engineer, GPUaaS
Design and maintain GPU clusters for AI/HPC workloads, automate provisioning, and optimize performance using Kubernetes, Slurm, and NVIDIA tools.
Devops Engineer (Strong in Java & Python)
Build and automate CI/CD pipelines for deploying and managing AI models (LLMs, agents) in a secure banking environment using Kubernetes, Docker, and cloud platforms.
DevOps Engineer
Design and maintain GPU clusters for AI/HPC workloads, automate provisioning, and optimize performance using Kubernetes, Slurm, and NVIDIA tools.