Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Research engineer optimizing large-scale generative AI training systems, focusing on performance, stability, and low-precision techniques for diffusion and multimodal models.
Develops and optimizes SIMD kernels for AI operators (e.g., softmax, layer norm) to run on next-gen hardware, focusing on performance and SDK usability for generative AI compute engines.
Build and optimize Cerebras' GPU-based AI inference stack, integrating vLLM, ROCm, and AMD GPUs to deliver ultra-low-latency, high-throughput LLM serving for production workloads.
Designs AI-driven systems that recursively optimize compute workloads and hardware, using agentic loops, performance profiling, and reinforcement learning to improve correctness, speed, and efficiency.
Neurosoft Bioelectronics is developing next-generation AI for decoding neural time series (high-density subdural ECoG LFPs). We seek an intern machine learning scientist with interest in sequence modelling, state-space…
Build and optimize cutting-edge large language models and foundation models using decentralized learning methods, focusing on post-training, distributed GPU training, and open-source deployment.
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology can change lives for the better. It can heal us, entertain us, and make us more connected, productive, and understanding of the world around us. And…
Develops and optimizes AI models for satellite and aerial imagery using generative techniques like Diffusion and GANs, focusing on super-resolution and image translation, and deploys them in on-premise GPU environments.
Designs and deploys AI/ML models for real-time network management using Python, PyTorch, and cloud platforms to power Extreme Networks' networking solutions.
Build and scale the AI/ML backbone for autonomous defense systems, including training pipelines, synthetic data generation, and real-time edge inference on embedded hardware.
Develops and optimizes Linux-based embedded systems for autonomous defense platforms, integrating high-bandwidth sensors and real-time AI inference on NVIDIA Jetson hardware.
Build and optimize AI model inference systems to serve millions of users with low-latency GPU workloads using frameworks like vLLM or Triton.
Architect next-gen GPU communication libraries (NCCL, NVSHMEM, UCX) to scale deep learning and HPC workloads across thousands of GPUs using C/C++, CUDA, and high-speed interconnects.
Own full-stack OS enablement for NVIDIA’s DGX Station AI workstation, ensuring seamless Windows and Linux support from firmware through AI application validation.
Build high-performance C++/CUDA storage libraries and algorithms to optimize GPU IO for AI workloads, working with Linux kernels, NVMe, and vector databases.
Build and scale the AI infrastructure that powers CrowdStrike’s LLM-driven security products, including GPU clusters, model-serving pipelines, and MLOps tooling.
Lead a team building next-gen simulation software in C++ for multi-physics systems, optimizing GPU-accelerated solvers and numerical algorithms used by automotive, aerospace and EV companies.
Sr. Manager, Future Computing Prototyping Studio Description - The Opportunity HP is creating a Shanghai-based Future Computing Prototyping Studio to make the next generation of personal computing tangible on real HP…
Lead a team building and optimizing NVIDIA’s open-source deep learning inference frameworks (e.g., SGLang, vLLM) to deploy LLMs and generative AI efficiently on GPUs from datacenter to edge.
We couldn't check your fit for this role — add a CV to your profile to see it next time.