Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Inference runtime engineer working at the core of vLLM, optimizing LLM and diffusion model serving across diverse hardware and model architectures (MoE, multimodal, agentic). Day-to-day is systems/ML engineering in Python and PyTorch to make AI inference cheaper and faster.
The Member of Technical Staff will build and maintain the cloud orchestration infrastructure for vLLM, focusing on cluster management, deployment automation, and observability for large-scale AI inference. The role requires expertise in Kubernetes, infrastructure-as-code, and GPU cluster management.
Hands-on cluster administration engineer owning and operating high-performance GPU/HPC compute infrastructure for an AI inference company, leveraging Linux, SLURM/Kubernetes, and automation tooling.
Site Reliability Engineer focused on ensuring vLLM’s AI inference engine operates reliably, scalably, and with minimal downtime by designing resilient systems, improving observability, and driving incident response improvements.
Design the visual identity and developer-facing interfaces for vLLM, an AI inference engine, from brand assets to product UIs and launch campaigns.
Overview Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at…
Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the…
We couldn't check your fit for this role — add a CV to your profile to see it next time.