Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and operate scalable GPU-powered infrastructure for real-time AI model inference using frameworks like vLLM or Triton, ensuring high availability and cost efficiency.
Build and optimize low-level AI kernels (GEMM, attention, quantization) for Meta’s custom AI chips, shaping hardware-software co-design to maximize performance for billions of users.
We’re hiring a hands-on Computer Vision Engineer to build and improve sports video intelligence models—detection, tracking, pose, event understanding, and multi-view reasoning. You’ll spend most of your time on CV…
Field Service Engineer Location: Swindon/Oxford - Field-Based Salary: £39,353 + Up to £1,000 Onboarding Bonus + Performance Bonus Company: NoteMachine (A Brink's Company) Join NoteMachine and Help Power the…
Join a world-class data science team at JPMorgan Chase and help shape the future of our Chief Administrative Office. As a leader in applied AI and machine learning, you’ll have the opportunity to work on high-impact…
RELOCATION ASSISTANCE: Relocation assistance may be available CLEARANCE REQUIRED FOR START: Yes CLEARANCE TYPE: Secret TRAVEL: Yes, 10% of the Time Description At Northrop Grumman, our employees have incredible…
Own the AI inference performance roadmap at NVIDIA, turning deep optimization techniques into products that improve latency, efficiency, and cost per token across the inference stack for LLM deployments.
Build and optimize scalable generative AI inference pipelines and APIs for Adobe’s Firefly suite, integrating models like diffusion transformers into Photoshop, Illustrator, and other products.
Build and own the shared AI platform that trains and serves Adobe’s generative AI products at global scale, focusing on GPU fleet utilization, model serving, and distributed systems for low-latency inference.
Memwize is a VC-backed, stealth-mode startup building rack-level AI inference systems . We are seeking a Senior/Principal Machine Learning Compiler Developer to build compiler capabilities from PyTorch through Triton…
Optimize and port AI inference/training kernels across NVIDIA, AMD, TPUs, and emerging accelerators, designing portable abstractions for SGLang and Miles.
Develops high-performance kernels, compilers, and communication libraries to optimize AI workloads on GPU clusters, focusing on low-latency and memory efficiency.
Engineer and optimize GPU clusters for AI training and inference, build scheduling/orchestration systems, and improve performance, cost, and reliability of AI infrastructure.
Optimize AI model inference for low-latency production deployment, manage Kubernetes-based ML infrastructure, and collaborate with data-science teams to scale and automate systems handling thousands of concurrent requests.
Обязанности: Сформировать дорожную карту ML‑инициатив для трейдинга и смежных функций (ценообразование, прогнозы спроса/предложения, оптимизация логистики, риск‑метрики); Разрабатывать end-to-end LLM-приложения : от…
Lead the design and deployment of large-scale generative AI voice and speech systems, building reusable ML infrastructure and guiding cross-functional teams to integrate cutting-edge AI into products.
Build and maintain ML infrastructure to deploy and scale AI models in production, using Python, Kubernetes, and Ray. Own platform services that enable data scientists to move models from experimentation to live systems.
Build and scale a high-performance search platform using OpenSearch, Python/FastAPI, and vector search (embeddings, ANN/HNSW) to power GigaChat and LLM queries with low-latency, relevance tuning, and hybrid search.
We couldn't check your fit for this role — add a CV to your profile to see it next time.