Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build the full-stack infrastructure for Sciforium’s multimodal AI models, including a low-latency chat UI, developer console, billing, and backend APIs using TypeScript, Python, and C++.
Builds and optimizes large-scale AI infrastructure, including Kubernetes clusters, RDMA networking, and GPU orchestration to improve efficiency and scalability of AI training/inference systems.
Build and run large-scale, cloud-native systems that power Adobe’s AI features, including ML inference infrastructure, container orchestration, and automated patch management across AWS, Azure, and GCP.
Lead the architecture and engineering of a bank-grade AI & Agentic Platform, designing agentic runtimes, LLM gateways, identity layers, and cloud-native infrastructure to support secure, scalable agent workflows across the enterprise.
Optimize LLM inference performance on Intel GPUs by profiling bottlenecks, writing custom kernels, and contributing to open-source serving frameworks like vLLM and SGLang.
Builds full-stack web apps in Python (Flask/Django) with React/TypeScript front ends, integrating LLM capabilities via RAG and processing large datasets with PostgreSQL, ArangoDB, and vector search.
Зарплата: 150000–250000 RUR (до вычета налогов) MCN Telecom - один из ведущих операторов фиксированной связи (в 67 городах РФ), виртуальный мобильный оператор (MVNO), разработчик программных продуктов. Мы развиваем…
Build and deploy AI inference systems with customers, writing production code and debugging across the full stack from requests to kernel dispatch.
Optimize and deploy large-scale on-prem LLM inference systems using NVIDIA H200 GPUs, vLLM, TensorRT-LLM, and OpenShift AI for enterprise private GenAI environments.
Build and operate scalable ML inference platforms for an AI-native cloud startup, designing GPU-powered serving systems, deployment pipelines, and observability for real-time AI applications.
Build and operate scalable ML inference platforms using vLLM/TGI/Triton to serve AI models with low latency and high GPU efficiency for a next-gen cloud startup.
About Inflection AI Inflection AI is a Public Benefit Corporation empowering people with human-centered, emotionally intelligent AI. We’re shaping the future of AI by combining emotional intelligence (EQ) and raw…
Build and deploy enterprise LLM applications, RAG systems, and AI agents using open-source models (DeepSeek, Qwen, Kimi) and frameworks like LangChain and vLLM.
Build and own AI-powered document processing systems: deploy open-source LLMs, design REST APIs, and integrate AI features into web apps with streaming responses and RAG pipelines.
Engineer AI/LLM inference on GPU clusters: benchmark, tune, and optimize model serving with vLLM, Triton, or TensorRT-LLM to hit latency, throughput, and memory targets.
Own and scale the reliability, performance, and cost of ManyChat's AI infrastructure, including LLM inference services and AI Gateway, while shaping standards for the company's AI platform.
Build and scale SentinelOne’s AI Gateway infrastructure (Kong-based) to route, secure, and monitor AI coding assistant traffic, while operating self-hosted LLM stacks and driving reliability across Kubernetes and CI/CD systems.
Build and scale SentinelOne’s AI Gateway infrastructure (Kong-based) to route, secure, and monitor AI coding assistant traffic, while operating self-hosted LLM stacks and driving reliability across Kubernetes and CI/CD tooling.
WHO WE ARE 🌍 Creating content that resonates is great — turning that attention into growth is even better. That's what Manychat does. Our AI-powered automations help creators and brands engage with audiences…
Salary: $150 – $170 per hour About the Role The Domain Architect - AI Storage acts as the primary technical authority for the physical and logical lifecycle of high-performance data platforms across diverse client…
We couldn't check your fit for this role — add a CV to your profile to see it next time.