Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and optimize NVIDIA’s deep-learning inference stack (TensorRT, TensorRT-LLM) by profiling models, writing GPU kernels, and integrating OSS frameworks to maximize GenAI performance across datacenter and edge GPUs.
Build and optimize CUDA features for AI frameworks like PyTorch and TRT-LLM, improving multi-GPU performance and distributed runtime for training and inference workloads.
Build and optimize GPU-powered inference systems for large language models, improving serving efficiency and performance through kernel tuning and novel algorithms.
Design and deploy AI/ML solutions on cloud GPU platforms, collaborating with customers to integrate NVIDIA’s full stack of hardware and software technologies.
Build and scale FastAPI-based REST services, set up CI/CD, and deploy ML models in a high-load ecommerce environment using Python, Kubernetes, and MLOps tooling.
Architect and own LSports’ GCP-based cloud platform and its agentic AI layer, designing scalable, secure systems for real-time sports data and LLM-powered automation.
Senior ML engineer building scalable GenAI inference pipelines and APIs that power Adobe’s Firefly, Photoshop, and other creative tools using PyTorch, CUDA, and diffusion models.
Build and maintain the ML platform that deploys, monitors, and scales AI models powering Hadrian’s autonomous factories, using MLflow, Dagster, EKS, and FastAPI.
Builds and optimizes large-scale AI infrastructure, including Kubernetes clusters, RDMA networking, and GPU orchestration to improve efficiency and scalability of AI training/inference systems.
Lead the architecture and engineering of a bank-grade AI & Agentic Platform, designing agentic runtimes, LLM gateways, identity layers, and cloud-native infrastructure to support secure, scalable agent workflows across the enterprise.
Optimize LLM inference performance on Intel GPUs by profiling bottlenecks, writing custom kernels, and contributing to open-source serving frameworks like vLLM and SGLang.
Lead end-to-end system validation for the Triton Program, designing and executing automated/manual tests to verify functional requirements and improve mission readiness.
Designs and tests hardware systems for aerospace/defense using CATIA V5 to model ground equipment, cables, and racks, and releases drawings via Teamcenter PLM.
NVIDIA’s accelerated computing platform relies on continuous performance excellence at every stage of development. We are seeking an outstanding Performance Analysis Manager to lead an engineering team responsible for…
Build and maintain the ML platform that deploys, monitors, and scales AI models for Hadrian’s autonomous factories, using MLflow, Dagster, EKS, and FastAPI.
Build and operate scalable ML inference platforms for an AI-native cloud startup, designing GPU-powered serving systems, deployment pipelines, and observability for real-time AI applications.
Build and operate scalable ML inference platforms using vLLM/TGI/Triton to serve AI models with low latency and high GPU efficiency for a next-gen cloud startup.
Lead the design, development, and deployment of recommender systems and personalization models for Penguin Random House’s digital platforms to improve book discovery and customer engagement.
Engineer AI/LLM inference on GPU clusters: benchmark, tune, and optimize model serving with vLLM, Triton, or TensorRT-LLM to hit latency, throughput, and memory targets.
Own and scale the reliability, performance, and cost of ManyChat's AI infrastructure, including LLM inference services and AI Gateway, while shaping standards for the company's AI platform.
We couldn't check your fit for this role — add a CV to your profile to see it next time.