Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and scale the enterprise DevOps operating model for GXO’s Agentic AI Platform on GCP and Kubernetes, owning secure CI/CD, Terraform infrastructure, and AI model-serving pipelines.
Build and optimize open-source LLM inference engines for CXL-based memory offloading, integrating cryptographic acceleration and contributing upstream to AI infrastructure projects.
Build and optimize AI infrastructure to measure and maximize compute efficiency for training and inference workloads, ensuring accurate capacity planning and performance at scale.
Build and optimize a locally hosted AI platform, deploying and running GPU-powered models, tuning inference pipelines, and integrating AI into internal tools.
AI Application Engineer (Part-Time) Location: Argentina, Brazil, Peru, Colombia, Costa Rica, Mexico Department: Workana Premium Workplace: remote Employment Type: full Description Client: medxprts.ai Location: Remote…
Senior DevOps Engineer building and scaling cloud-native infrastructure, Kubernetes clusters, and AI serving platforms using IaC, GitOps, and observability tools.
Build and operate the ML infrastructure powering a global AI assistant, including training, deployment, inference, and observability systems in Python and PyTorch/JAX.
Design and run CI/CD, MLOps, and Kubernetes-based cloud infrastructure for a Bitcoin-mining and AI-cloud company, ensuring high availability and security across global datacenters.
Build and refine LLM models for fintech clients, using prompt engineering, OCR, NLP, and multimodal approaches to power industrial systems.
Builds and maintains Java/Spring Boot backend services and REST APIs, leveraging AI tools to accelerate development, code review, testing, and debugging in a collaborative engineering team.
Looking for an AI Engineer to Develop, deploy, and operate AI/LLM models across Clinets dual environment — GCP for public-cloud workloads, Humain sovereign cloud for classified data. Requirements Build and fine-tune…
Forward Deployed Engineer at DeepInfra designs and runs AI inference benchmarks, tunes deployments on cutting-edge hardware, and partners with sales to win enterprise deals.
Own the roadmap for WEKA’s Augmented Memory Grid, optimizing LLM inference performance by offloading KV-cache and integrating with engines like vLLM and NVIDIA Triton.
Job Title: MLOps Platform Developer Location: London Salary: Depending on experience Job Type: Full time, Permanent This is a rare opportunity to be first in the door to continue the development of our engineering…
Build and operate an LLM-powered content generation pipeline with safety guardrails, RAG grounding, and image generation, deployed on AWS serverless services.
Research and implement model quantization algorithms to optimize AI models for on-device deployment using PyTorch, ONNX, and related tools.
R&D intern on Nota’s AI team optimizing models for deployment on edge devices using PyTorch, ONNX, and frameworks like ExecuTorch/TensorRT.
Research and develop quantization, pruning, and inference optimizations for LLM/VLM and MoE models to run efficiently on GPUs and NPUs.
Builds data pipelines and inference pipelines for a local LLM-based clinical support agent, normalizing hospital EMR data and optimizing for on-premise latency/resource constraints.
Builds and maintains LLM training, fine-tuning, and inference pipelines on AWS using PyTorch, vLLM, and FastAPI to power a healthcare platform’s AI features.
We couldn't check your fit for this role — add a CV to your profile to see it next time.