Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
NVIDIA’s accelerated computing platform relies on continuous performance excellence at every stage of development. We are seeking an outstanding Performance Analysis Manager to lead an engineering team responsible for…
Build and operate scalable ML inference platforms for an AI-native cloud startup, designing GPU-powered serving systems, deployment pipelines, and observability for real-time AI applications.
Build and operate scalable ML inference platforms using vLLM/TGI/Triton to serve AI models with low latency and high GPU efficiency for a next-gen cloud startup.
Build and optimize next-gen ML models in C++/Python with CUDA and PyTorch for real-world AI applications.
Senior engineer building the backend infrastructure for an AI acceleration cloud, including Kubernetes clusters, high-performance storage, and GPU virtualization across global data centers.
Engineer AI/LLM inference on GPU clusters: benchmark, tune, and optimize model serving with vLLM, Triton, or TensorRT-LLM to hit latency, throughput, and memory targets.
Builds and optimizes real-time computer-vision pipelines using Python, C/C++, OpenCV, and GStreamer for defense-tech systems.
Lead a team building real-time radar and electronic-warfare signal-processing software in C++/CUDA, architecting DSP pipelines for deployable systems and validating performance via simulation and field tests.
Research and develop on-device generative and 3D spatial AI models for mobile platforms, optimizing models for edge deployment and shipping to product.
Salary: $150 – $170 per hour About the Role The Domain Architect - AI Compute acts as the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across diverse client…
Build and maintain the Kubernetes-based infrastructure and MLOps tooling that trains, validates, and deploys AI models for Intuitive Surgical’s medical devices, ensuring reproducible workflows and GPU cluster health.
Solution Architect designs enterprise-scale GenAI/LLM systems, including RAG pipelines, agent orchestration, and secure on-prem deployments, while guiding clients through AI transformation and regulatory compliance.
Build and operate a unified GPU compute platform for AI training and inference, hiding cloud complexity behind Kubernetes operators and schedulers that manage multi-cloud NVIDIA clusters.
Junior engineer builds and maintains ML/HPC infrastructure for trading systems, learning large-scale reliability and performance at Squarepoint Capital.
Build and optimize real-time graphics and ML inference pipelines for interactive visual apps, integrating super-resolution and denoising models into DX12/Vulkan renderers.
ML engineer builds and deploys local LLM pipelines, RAG systems, and NLP models entirely offline for a state-run transport data platform.
Optimize and port AI inference/training kernels (SGLang, Miles) across NVIDIA/AMD GPUs, TPUs, CPUs, and emerging accelerators to maximize performance on heterogeneous hardware.
Develop GPU-accelerated CFD simulation software with AI-assisted methods, designing solvers and APIs for modern engineering workflows.
Build and optimize GPU kernels and inference frameworks (e.g., vLLM) to accelerate large language model serving, integrating research into production-grade, open-source software.
Build and optimize on-device AI inference software for NVIDIA GPUs, focusing on low-latency, memory efficiency, and deployment on RTX/DGX systems.
We couldn't check your fit for this role — add a CV to your profile to see it next time.