Tech jobs
Job listings
Machine Learning Engineer — Inference Optimization
About the Role We’re looking for a Machine Learning Engineer to own and push the limits of model inference performance at scale . You’ll work at the intersection of research and production—turning cutting-edge models…

Member of Technical Staff, Cloud Orchestration
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the…

Member of Technical Staff, Kernel Engineering
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the…

Member of Technical Staff, Performance and Scale
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the…
Software Engineer, Inference
About Luma AI Luma's mission is to build multimodal AI to expand human imagination and capabilities. We believe that multimodality is critical for intelligence. To go beyond language models and build more aware,…
Senior Software Engineer - Model Performance
Help us make inference blazingly fast. If you love squeezing every last drop of performance out of GPUs, diving deep into CUDA kernels, and turning optimization techniques into production systems, we'd love to meet…
Senior Product Manager, AI Models
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based…
Researcher - LLMs
Huawei Canada has an immediate 12-month contract opening for a Researcher. About the team: Founded in 2012, the Noah’s Ark lab has evolved into a prominent research organization with notable achievements in academia…
Principal AI Platform Engineer (CA)
The Team This team will serve as the product owner for GenAI capabilities within PointClickCare, working closely with other engineering teams across the organization to identify, build and support generative AI…
Principal AI Platform Engineer (US)
The Team This team will serve as the product owner for GenAI capabilities within PointClickCare, working closely with other engineering teams across the organization to identify, build and support generative AI…
CVP of Applied AI FDE
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a…
Développeur Full stack
En tant que Développeur Full Stack vous participez à la conception, au développement et à la maintenance des applications internes et clientes mettant en oeuvre des modules d’IA Au sein de notre Direction Technique et…

Machine Learning Engineer
About Osmosis At Osmosis, we help companies use cutting-edge reinforcement learning techniques to fine-tune open-source language models that beat foundation models on performance, latency, and cost. We’ve raised $7M in…
Senior Deep Learning Algorithm Engineer
We are now looking for a Senior DL Algorithms Engineer! NVIDIA is seeking senior engineers who are mindful of performance analysis and optimization to help us squeeze every last clock cycle out of Deep Learning…
Backend Lead - Python - Dallas, TX
We are looking for a Backend Lead to architect and build the robust infrastructure required to power our Agentic AI products. You will lead the development of the "brain" of our platform—the layer where LLMs…
Senior Inference Technical Product Marketing Manager - Accelerated Computing
We are looking for a Senior Technical Product Marketing Manager. This role will be located in our rapidly growing data center business and pivotal in our inference marketing. You will be focused on working with…
Principal Software Engineer – Large-Scale LLM Memory and Storage Systems
NVIDIA Dynamo is a high-throughput, low-latency inference framework for serving generative AI and reasoning models across multi-node distributed environments. Built in Rust for performance and Python for extensibility,…
Senior Deep Learning Software Engineer, Inference
NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize the GPU-accelerated software that powers today’s…
Staff MLOps Engineer, LLMOps
Build a Safer World. TRM Labs provides AI-powered intelligence solutions that help public and private sector agencies investigate and disrupt crime. TRM's platforms enable investigators to trace illicit activity,…