Tech jobs
Job listings
Staff Machine Learning Infrastructure Engineer, Embedding Platform
Build and scale Reddit’s next-gen ML embedding platform, designing distributed training and low-latency serving systems that power personalized recommendations across the site.
Senior Machine Learning Engineer – Predictive World Model
Build predictive world models using generative AI and deep learning to simulate how scenes evolve, training autonomous driving and robotics policies from multimodal sensor data.
Member of Technical Staff - Training Platform
Build and operate a hosted training platform for AI models, including Kubernetes orchestration, GPU scheduling, observability, and a Next.js/React UI.
Member of Technical Staff - Compute Platform
Builds and maintains the platform that schedules, monitors, and runs AI workloads, using Python, Rust, Kubernetes, and cloud infrastructure.
AI Research Engineer - Datadog AI Research (DAIR)
Research Engineer at Datadog’s AI lab in Paris builds and scales multimodal AI systems for cloud observability, training world models and autonomous agents using PyTorch, Ray, and GPU clusters.
Manager I, Engineering - AI Platform - Training & Serving
The AI platform is responsible for all AI infrastructure across Datadog. Our mission is to provide tools and platforms that enable data scientists and engineers to conduct large-scale training and inference with ease.…
Product Manager II - Model Lab
Build and launch Datadog’s new experiment tracking platform for AI/ML teams, integrating metrics, datasets, and model lineage into a unified observability product.
Senior System Administrator
What You'll Do: Looking for an opportunity to apply your technical expertise to a mission that matters, this is it. As the Senior System Administrator , you will support the systems and technologies that power the…
Member of Technical Staff — Inference-Multi-Hardware
Optimize and port AI inference/training kernels across NVIDIA, AMD, TPUs, and emerging accelerators, designing portable abstractions for SGLang and Miles.
Member of Technical Staff — Inference-Multimodal & Diffusion
Research and engineer next-generation diffusion and flow-based models for image, video, and multimodal generation, scaling training and deployment on GPUs/TPUs.
Senior ML Engineer (AI Research, Physical AI)
Build and train vision-language-action models and reinforcement-learning policies for real-world robots, using JAX and distributed training to prototype sim-to-real transfer and dexterous manipulation.
AI/ML Specialist Solutions Architect
Design and deploy scalable AI solutions for enterprise customers, advising on MLOps, distributed training, and GPU orchestration using PyTorch, Kubernetes, and related tools.
Member of Technical Staff — Inference-Kernel, Compiler & Communication
Develops high-performance kernels, compilers, and communication libraries to optimize AI workloads on GPU clusters, focusing on low-latency and memory efficiency.
Machine Learning Performance Engineer - Offboard Training & Inference
Optimize large-scale ML training and batch inference for autonomy workloads, profiling GPU clusters to maximize throughput and reduce cost per unit of data processed.
Staff Machine Learning Engineer - Moloco Commerce Media
Build and productionize ML models for ad ranking, recommendations, and search in Moloco’s Commerce Media platform, driving performance for retail media customers.
Senior Software Engineer, Agent Platform (AI for the Planet)
Build and operate an AI agent platform for conservation and climate teams, designing APIs, SDKs, and observability tools while shipping real-world agents for users like NASA JPL and Global Mangrove Watch.
Senior ML Engineer (AI Research, Physical AI)
Build and train vision-language-action models and reinforcement-learning policies for real-world robots, using JAX and distributed training to prototype sim-to-real transfer and dexterous manipulation.
Cluster Network Engineering Lead
Lead the design and operation of high-performance AI cluster networks, focusing on InfiniBand/RoCE fabrics, RDMA performance, and fabric reliability for large-scale training and inference workloads.
Senior Principal AI Engineer
Lead the design and optimization of large-scale distributed AI training systems, improving performance and scalability for advanced neural networks across GPU clusters.