Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build AI agents, fine-tune LLMs, and deploy RAG systems for internal security tools using PyTorch, LangChain, and vLLM.
Build and maintain automated QA flows and benchmarks for AI frameworks like PyTorch and vLLM, validating performance across hardware and integrating tests into CI/CD pipelines.
Develop and optimize AI frameworks like vLLM and PyTorch, focusing on distributed algorithms and performance tuning for deep learning models across hardware backends.
Develops and optimizes AI frameworks like SGLang, profiling distributed deep learning models to improve performance across hardware backends.
Build and maintain Python-based backend systems, AI workflows, and React UIs for Proofpoint’s cybersecurity platform, handling data capture, compliance, and LLM-powered investigations.
Own the AI inference performance roadmap at NVIDIA, turning deep optimization techniques into products that improve latency, efficiency, and cost per token across the inference stack for LLM deployments.
Build and maintain agentic infrastructure for aerospace systems, including local LLM inference, multi-agent systems, and production-grade compute environments to support rocket development and factory operations.
Build and maintain agentic infrastructure for aerospace systems, including local LLM inference, multi-agent systems, and production-grade compute environments to support rocket development and manufacturing.
Build and own the shared AI platform that trains and serves Adobe’s generative AI products at global scale, focusing on GPU fleet utilization, model serving, and distributed systems for low-latency inference.
Memwize is a VC-backed, stealth-mode startup building rack-level AI inference systems . We are seeking an experienced Senior/Principal Machine Learning Engineer to drive performance improvements in inference frameworks…
Lead AI inference optimization and deployment for autonomous-vehicle models, tuning low-latency pipelines on embedded and cloud GPUs with CUDA, TensorRT, and LLM serving stacks.
Build and operate a secure Kubernetes-based platform for AI engineering tools, implementing GitOps and observability while ensuring sovereignty, reliability, and auditability in a high-security environment.
Build and secure AWS/Azure cloud infrastructure using IaC, CI/CD automation, and Kubernetes hardening, embedding security-by-design across the SDLC and integrating Microsoft Defender for Cloud.
Optimize and port AI inference/training kernels across NVIDIA, AMD, TPUs, and emerging accelerators, designing portable abstractions for SGLang and Miles.
Build and maintain the ML platform that trains, deploys, and monitors AI/ML models across Wealthsimple’s investing, crypto, and other financial products.
Engineer and optimize GPU clusters for AI training and inference, build scheduling/orchestration systems, and improve performance, cost, and reliability of AI infrastructure.
Build and maintain AI inference infrastructure using Python, Kubernetes, and Helm to deploy and optimize LLM services for internal and customer use.
Optimize AI model inference for low-latency production deployment, manage Kubernetes-based ML infrastructure, and collaborate with data-science teams to scale and automate systems handling thousands of concurrent requests.
Обязанности: •Разработка ИИ агентов и мультиагентной платформы: логика агентов, взаимодействие, оркестрация •Интеграция с LLM: подключение моделей, промпты, обработка ответов API-разработка: REST API для интеграции с…
We couldn't check your fit for this role — add a CV to your profile to see it next time.