Tech jobs
Job listings
Staff/Principal DevOps Engineer, AI Inference
Designs and runs GPU-powered infrastructure for serving machine learning models at scale, optimizing low-latency inference on Kubernetes and AWS accelerators.
Staff Software Engineer, GPU Inference
Build and optimize Cerebras' GPU-based AI inference stack, integrating vLLM, ROCm, and AMD GPUs to deliver ultra-low-latency, high-throughput LLM serving for production workloads.
AI Engineer, Recursive Self-Improvement for Compute
Designs AI-driven systems that recursively optimize compute workloads and hardware, using agentic loops, performance profiling, and reinforcement learning to improve correctness, speed, and efficiency.
Founding ML Engineer in the Flower Frontier Model Team (all seniority levels welcome) [UK, Germany, Global]
Build and optimize cutting-edge large language models and foundation models using decentralized learning methods, focusing on post-training, distributed GPU training, and open-source deployment.
ML Engineer
Develops and optimizes AI models for satellite and aerial imagery using generative techniques like Diffusion and GANs, focusing on super-resolution and image translation, and deploys them in on-premise GPU environments.
Machine Learning Engineer
Build and scale the AI/ML backbone for autonomous defense systems, including training pipelines, synthetic data generation, and real-time edge inference on embedded hardware.
Embedded Autonomy Engineer
Develops and optimizes Linux-based embedded systems for autonomous defense platforms, integrating high-bandwidth sensors and real-time AI inference on NVIDIA Jetson hardware.
Software Engineer (Model Inference)
Build and optimize AI model inference systems to serve millions of users with low-latency GPU workloads using frameworks like vLLM or Triton.
Senior Software Architect - Deep Learning and HPC Communications
Architect next-gen GPU communication libraries (NCCL, NVSHMEM, UCX) to scale deep learning and HPC workloads across thousands of GPUs using C/C++, CUDA, and high-speed interconnects.
Senior Systems Software Engineer, Windows and Linux Enablement - DGX Station
Own full-stack OS enablement for NVIDIA’s DGX Station AI workstation, ensuring seamless Windows and Linux support from firmware through AI application validation.
Senior Software Engineer, AI Storage
Build high-performance C++/CUDA storage libraries and algorithms to optimize GPU IO for AI workloads, working with Linux kernels, NVMe, and vector databases.
Sr. AI Infrastructure Engineer, LLM/AI Platforms (Remote)
Build and scale the AI infrastructure that powers CrowdStrike’s LLM-driven security products, including GPU clusters, model-serving pipelines, and MLOps tooling.

Lead Software Engineer
Lead a team building next-gen simulation software in C++ for multi-physics systems, optimizing GPU-accelerated solvers and numerical algorithms used by automotive, aerospace and EV companies.
Sr. Manager, Future Computing Prototyping Studio
Sr. Manager, Future Computing Prototyping Studio Description - The Opportunity HP is creating a Shanghai-based Future Computing Prototyping Studio to make the next generation of personal computing tangible on real HP…
Engineering Manager, Deep Learning Inference
Lead a team building and optimizing NVIDIA’s open-source deep learning inference frameworks (e.g., SGLang, vLLM) to deploy LLMs and generative AI efficiently on GPUs from datacenter to edge.
Sr. Manager, Future Computing Prototyping Studio
Sr. Manager, Future Computing Prototyping Studio Description - The Opportunity HP is creating a Shanghai-based Future Computing Prototyping Studio to make the next generation of personal computing tangible on real HP…
Senior Machine Learning Engineer I, Physical Sciences
Senior ML engineer builds and deploys scalable AI pipelines for materials, chemistry and physical sciences, using PyTorch, cloud infra and production-grade MLOps.
Staff Software Engineer - AI Platform
Builds and optimizes LinkedIn’s large-scale AI training, feature engineering, and serving infrastructure for recommendation, LLM, and computer vision models using distributed systems, GPU clusters, and open-source frameworks.