Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Senior Software Engineering Manager – KV Cache Platform Department: Top Level - By Department Location: Santa Clara Employment Type: FullTime DDN is seeking a Senior Software Engineering Manager to lead the engineering…
Support and architect AI platforms for DDN’s Hyperpod, diagnosing issues across NVIDIA AI Enterprise, vector databases, GPUs, Kubernetes, and high-performance storage/networking.
Builds and optimizes large-scale AI inference systems for frontier models, focusing on performance, latency, and cost across thousands of GPUs.
Principal Architect leads the design of a GPU-powered simulation platform for AI data centers, enabling real-time power orchestration and validation of energy-aware compute systems.
Performance Engineer at RadixArk in Palo Alto optimizes LLM inference and training systems for latency, throughput, and cost efficiency across production workloads using SGLang, Miles, and GPU/TPU infrastructure.
Optimize and accelerate LLM inference and training systems like SGLang and Miles by profiling GPU performance, writing custom kernels, and enabling new models on modern hardware.
Build and ship production-grade AI systems—LLMs, agents, retrieval, and evals—from prototype to scalable deployment for high-stakes decision-making.
Build and scale the ML platform powering PubMatic’s adtech stack, designing pipelines for petabyte-scale data, GPU-accelerated training/inference, and RAG systems to troubleshoot bid streams and deliver AI-driven insights.
Build and evaluate MLOps pipelines for cutting-edge GenAI models using JAX, PyTorch, and GPU kernels (Pallas/Triton) to improve training data quality.
Build and operate a secure, scalable platform to host large language models, integrating GPU clusters, MLOps pipelines, and observability while optimizing inference performance and tenant isolation.
Mission 4Minds is an enterprise AI fine-tuning platform that transforms how organizations build and operate private, domain-specific AI. Unlike static systems, 4Minds’s AI platform learns continuously from live data in…
Build and optimize production-grade inference services for Apple’s large language, vision, and speech models, ensuring low-latency performance across products like Siri and Photos.
Build and deploy India’s open healthcare foundation model: train a 30B-parameter MoE on 500B+ Indic-language tokens, then ship fast inference and robust evals for doctors and government use.
Build and optimize distributed AI training systems, profiling bottlenecks in PyTorch stacks and low-level GPU kernels to speed up model convergence.
Build and optimize a compiler stack for robotics simulation and AI training, using LLVM, JIT, and GPU codegen to maximize performance.
Design the hardware and software stack that powers Lumai’s optical AI accelerator, ensuring it meets the compute and memory demands of modern transformer models and other AI workloads.
Build and scale generative and predictive ML models for cellular behavior using PyTorch and distributed training, bridging research prototypes to production-grade systems in a TechBio company.
Build and deploy AI agents for fintech workflows (risk, fraud, payments) and the platform that powers them, including orchestration, tooling, and safety guardrails.
Principal AI/ML Architect designs and advises on production ML systems, MLOps/LLMOps pipelines, and GenAI architectures on AWS for enterprise clients, translating technical depth into business value.
Builds and secures cloud infrastructure and AI pipelines for a YC-backed fintech startup, focusing on AWS, Kubernetes, and SOC2 compliance while optimizing GPU workloads for real-time voice AI.
We couldn't check your fit for this role — add a CV to your profile to see it next time.