Tech jobs
Job listings
AI Tech Architect
Design and build cloud-native AI infrastructure, MLOps pipelines, and multi-modal feature platforms for fintech clients, leveraging AWS, GCP, Azure, and AliCloud.
Senior MLOps Engineer
Senior MLOps Engineer builds and optimizes ML pipelines, tracks experiments, and deploys models using tools like MLflow and PyTorch FSDP.
ML Ops Engineer
Build and scale the MLOps/LLMOps infrastructure for Anaplan’s AI-infused scenario-planning platform, automating model training, deployment, and GPU-optimized inference in Kubernetes.
Machine Learning Engineer - ML Training Platform
Build and optimize a distributed ML training substrate that trains large models across many low-bandwidth nodes using model parallelism and P2P networking.
AI Platform and MLOps Engineer
Build and maintain an AI/MLOps platform to deploy large language models and AI agents, optimize distributed training, and manage GPU clusters on public clouds.
Lead Machine Learning Engineer (Foundation Models)
Lead the development of Grab’s proprietary foundation models and generative recommendation systems, scaling distributed training and deploying AI solutions for millions of users.
Senior Associate, Generative AI, Data and Analytics, Advisory
Design and optimize ML pipelines, fine-tune LLMs, and deploy scalable LLM services using frameworks like MLflow, vLLM, and Kubernetes.
Senior Associate, Generative AI, Data and Analytics, Advisory
Design and deploy ML pipelines, fine-tune LLMs, and orchestrate GPU-based serving for enterprise GenAI solutions using frameworks like PyTorch, LangChain, and Kubernetes.
ML Infrastructure Engineer
Design and operate GPU clusters, distributed training frameworks, and scheduling systems to power large-scale AI workloads with a focus on reliability, efficiency, and cost control.
Senior Machine Learning Engineer – Predictive World Model
Build predictive world models using generative AI and deep learning to simulate how scenes evolve, training autonomous driving and robotics policies from multimodal sensor data.
AI Research Engineer - Datadog AI Research (DAIR)
Research Engineer at Datadog’s AI lab in Paris builds and scales multimodal AI systems for cloud observability, training world models and autonomous agents using PyTorch, Ray, and GPU clusters.
Machine Learning Performance Engineer - Offboard Training & Inference
Optimize large-scale ML training and batch inference for autonomy workloads, profiling GPU clusters to maximize throughput and reduce cost per unit of data processed.
Senior Principal AI Engineer
Lead the design and optimization of large-scale distributed AI training systems, improving performance and scalability for advanced neural networks across GPU clusters.
AI Enterprise Architect
Design and lead end-to-end AI infrastructure, from GPU clusters and high-performance networking to MLOps and governance frameworks, ensuring scalable, compliant, and production-ready AI systems.
Senior Site Reliability Engineer
Senior SRE builds and maintains Voleon’s AI research compute clusters, ensuring 24/7 uptime and performance for ML workloads across on-prem and cloud using IaC, observability stacks, and SRE practices.
Machine Learning Infrastructure Engineer
Build and scale ML infrastructure for brain-computer interface R&D, including distributed training pipelines and large-scale data platforms to support neuroscientific modeling and neural decoding.
AI Infrastructure Engineer
Design and operate GPU clusters, distributed training frameworks, and scheduling systems to power large-scale AI workloads with a focus on reliability, efficiency, and cost control.
Member of Technical Staff — Product
Builds and maintains developer-facing tools (APIs, SDKs, dashboards) on top of AI inference/training infrastructure like SGLang and Miles, collaborating with product and research teams to create intuitive interfaces.