Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Research and develop post-training methods (RLHF, SFT, PEFT, reward-based optimization) for Enchant, Iambic's large multimodal transformer model used in drug discovery, using Python and PyTorch at scale.
Research and develop post-training methods for a large multimodal transformer model (Enchant) in drug discovery, focusing on fine-tuning, reinforcement learning, and evaluation frameworks to advance AI-driven therapeutic development.
Space is a warfighting domain. True Anomaly seeks those with the talent and ambition to build the technology that secures it. OUR MISSION True Anomaly delivers decisive capabilities for space superiority. We build…
Developer Advocate at Fireworks AI builds and promotes hands-on technical content (tutorials, demos, guides) to help developers deploy and fine-tune AI models in production. Combines engineering (coding, model deployment) with public technical communication (blogs, talks, community engagement) to shape Fireworks as the go-to platform for AI infrastructure.
Senior Forward Deployed Software Engineer on ServiceNow's Applied AI team, building and deploying production LLM-powered applications end-to-end for strategic enterprise customers in London, spanning backend services, orchestration pipelines, and front-end integrations.
Lead the product strategy for Dialpad's AI models (SLMs, ASR, and inference infrastructure), balancing model quality, latency, and cost while owning the full lifecycle from data to production.
Build and operate large language model serving infrastructure at scale using Python, Kubernetes, and cloud platforms, applying site reliability engineering practices to AI platforms at J.P. Morgan.
Build and scale dunnhumby’s Enterprise AI Platform, designing production-grade AI systems including RAG, agentic workflows, and LLM-powered services using Python, LangChain, and cloud-native tools.
Own the end-to-end ML model lifecycle—training data, fine-tuning, evaluation, and release—for an on-premise LLM system that runs on customer hardware inside OPSWAT's MetaDefender Core cybersecurity platform. Core tech: Python, LLM fine-tuning (LoRA/QLoRA), inference serving (vLLM, Ollama), and integration with a Rust production service.
AI Consultant providing technical leadership for building and deploying scalable ML/LLM/SLM applications—including RAG pipelines, custom agents, multimodal systems, and cloud/MLOps infrastructure—on a part-time remote contract basis.
Principal engineer optimizing large-language-model inference at a major bank, designing quantization, speculative decoding, and benchmarking strategies to cut cost and latency for production AI workloads.
Develops and deploys generative AI and deep learning solutions for Citi’s Treasury & Trade Services, extracting insights from unstructured data (emails, call transcripts) to drive client experience and revenue growth. Core focus: end-to-end AI product delivery, LLM fine-tuning, RAG systems, and MLOps for production-grade GenAI.
The ML Platform Engineer will design, build, and maintain high-performance inference platforms for serving large machine learning models in production. The role focuses on systems engineering tasks such as request routing, GPU utilization, autoscaling, and observability for LLMs and other AI workloads.
Build high-performance C++ SDKs for computer vision, deploy and optimize ML models across diverse hardware (Qualcomm SoCs, Intel CPUs, NVIDIA GPUs), and maintain Python-based CI/CD automation pipelines at Zebra Technologies in London.
Design, optimize, and deploy machine learning models on resource-constrained edge devices using model compression, quantization, and frameworks like TensorFlow Lite, ONNX Runtime, and Core ML.
PhD researcher (EDB-IPP programme) focused on LLM model compression and acceleration—developing quantization, pruning, knowledge distillation, and inference optimization techniques using PyTorch for large language, multimodal, and diffusion models.
Builds and optimizes ML infrastructure for Nuro’s autonomous vehicle fleet, focusing on model compression, deployment, and performance improvements (e.g., quantization, distillation) to enhance real-world autonomy.
Builds and optimizes ML infrastructure for Nuro’s autonomous vehicle fleet, focusing on model pipelines, compilers (e.g., FTL), and deployment of optimized models for safe on-road navigation.
Build and operate the AI agent infrastructure that automates engineering tasks at Nuro, including closed-loop evaluation, agent platform, and autoresearch systems for autonomous driving.
Build and deploy AI systems for public safety, including computer vision, speech recognition, NLP, and generative AI, across cloud and edge devices.
We couldn't check your fit for this role — add a CV to your profile to see it next time.