Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and deploy quantized large language models for autonomous-vehicle inference, focusing on PTQ/QAT, mixed-precision, and runtime integration with PyTorch and TensorRT-LLM.
Build and optimize open-source LLM inference engines for CXL-based memory offloading, integrating cryptographic acceleration and contributing upstream to AI infrastructure projects.
Owns the full ML lifecycle for physical AI, from sensor data pipelines to deploying optimized models on constrained hardware, collaborating with embedded teams to ensure reliability and performance in real-world devices.
Build low-level compilers, runtimes, and AI/DSP optimizations for next-gen hardware in C++/Python, working closely with hardware teams.
AI Application Engineer (Part-Time) Location: Argentina, Brazil, Peru, Colombia, Costa Rica, Mexico Department: Workana Premium Workplace: remote Employment Type: full Description Client: medxprts.ai Location: Remote…
Looking for an AI Engineer to Develop, deploy, and operate AI/LLM models across Clinets dual environment — GCP for public-cloud workloads, Humain sovereign cloud for classified data. Requirements Build and fine-tune…
Forward Deployed Engineer at DeepInfra designs and runs AI inference benchmarks, tunes deployments on cutting-edge hardware, and partners with sales to win enterprise deals.
Lead FPGA firmware architecture for a portable SIGINT/EW platform, migrating GPU DSP algorithms to FPGA and optimizing real-time RF signal processing for defense missions.
Own the roadmap for WEKA’s Augmented Memory Grid, optimizing LLM inference performance by offloading KV-cache and integrating with engines like vLLM and NVIDIA Triton.
Research and implement model quantization algorithms to optimize AI models for on-device deployment using PyTorch, ONNX, and related tools.
Research and develop quantization, pruning, and inference optimizations for LLM/VLM and MoE models to run efficiently on GPUs and NPUs.
Develops embedded C/C++ firmware for sensor-based AI systems, integrating ML models (PyTorch/TensorFlow) with DSP and quantization for consumer electronics and automotive.
Build, optimize, and deploy large language and multimodal models for industrial use, focusing on training, compression, RAG, and agent workflows.
Lead AI Engineer builds agentic AI systems, RAG pipelines, and fine-tuned language models while ensuring security and compliance for a banking-focused AI product.
Build and own Life360’s on-device ML platform for edge devices like Tile and Pet GPS trackers, integrating ML models into resource-constrained firmware while optimizing for power, memory, and latency.
You'll join a company where your work has visible impact from day one. You'll work directly with world-class engineers, scientists, aerospace experts and entrepreneurs building technologies that have never…
Build and maintain embedded firmware for telematics and video devices that power fleet-management insights, optimizing safety, cost, and sustainability for thousands of vehicles.
Optimizes and deploys large language and diffusion models for on-device Apple Intelligence features, focusing on model compression, distillation, and hardware-aware techniques.
Build and lead Mercor’s code-search systems: hybrid retrieval (dense embeddings + BM25) that routes coding tasks to the right models and turns natural-language questions into precise code retrieval at scale.
Build and deploy ML models end-to-end: train, optimize, and package solutions while setting up MLOps pipelines, monitoring, and data governance for production AI systems.
We couldn't check your fit for this role — add a CV to your profile to see it next time.