Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Senior DevOps/MLOps Engineer at VinFast (EV maker) who builds and runs the infrastructure behind its Agentic AI and VoiceAI systems: deploying LLM model serving on AWS EKS with GPU optimization (vLLM, Triton), plus IaC, CI/CD, auto-scaling and observability for ultra-low-latency production AI services.
Build and optimize LLM-based AI agents for VinFast's SLP Center in Hanoi — designing multi-agent orchestration, RAG pipelines, function calling, and evaluation frameworks, plus fine-tuning and inference optimization. Core stack: Python, LangChain/LangGraph, PyTorch, Hugging Face, vLLM, Docker, and cloud platforms.
VinFast is hiring a Foundation AI Engineer to research, pretrain, and post-train LLMs/SLMs powering its robot and virtual assistant products, including Vietnamese-language domain adaptation, RLHF/DPO alignment, distributed training on GPU clusters, and model distillation for on-device deployment.
AI Engineer at EV maker VinFast building an in-vehicle conversational AI assistant: fine-tuning and evaluating LLMs (SFT, LoRA, RLHF/DPO), building RAG and multi-agent orchestration, optimizing models for on-device edge inference, and running production MLOps/LLMOps. Core stack: Python, PyTorch, HuggingFace, vLLM, TensorRT/ONNX/TFLite.
Contractor role at RavenPack (Marbella, remote with EU timezone) doing hands-on optimization of the search & LLM stack: fine-tuning open-source LLMs (PEFT/LoRA, DPO), inference acceleration (quantization, vLLM, Triton), retrieval improvements, and reproducible delivery via SageMaker and Docker. Initial 3-6 month contract.
Senior MLOps Solutions Engineer on the Pure Solutions team who designs and automates end-to-end MLOps pipelines and GPU-accelerated AI/ML reference architectures, integrating Pure Storage platforms (FlashBlade, FlashArray, Portworx) with Kubeflow, MLflow, and Ray. Day to day: CI/CD-driven pipeline automation, Terraform/Ansible-based infrastructure, and optimizing LLM inference with tools like NVID
Qutwo, a European quantum-AI lab in Helsinki, hires senior full-stack ML engineers to turn DNN-compression ML pipelines into a secure, production SaaS product. Day to day: Python backend services, frontends, and owning deployment on containers/Kubernetes with CI/CD, observability, and security.
About DevRev At DevRev, we're building the future of work with Computer – your AI teammate. Unlike traditional tools, Computer unifies all your data sources, tools, and workflows into a single AI-ready platform,…
When you join Verizon You want more out of a career. A place to share your ideas freely — even if they’re daring or different. Where the true you can learn, grow, and thrive. At Verizon, we power and empower how people…
NVIDIA’s accelerated computing platform is enabling the generational improvements in large language models, while the scale and complexity of these models are creating new challenges in computational efficiency. We are…
A hybrid technical/commercial role owning enterprise customer relationships for Fireworks AI's inference platform — from scoping pilots and onboarding through production deployment, technical escalations, QBRs, and renewals/expansions. Core tech: APIs, AWS/GCP/Azure, and production LLM work (prompting, fine-tuning, RAG).
Develops and optimizes AI models (LLMs/vision) for custom hardware, tuning performance and accuracy across software, compiler, and hardware layers.
Software Engineer on the AI Models/System Bring-Up team responsible for porting, validating, and optimizing AI models on Tenstorrent platforms. Core technologies include PyTorch, TensorFlow, JAX, Python, C++, and Linux environments.
A Machine Learning Research Engineer working on LLM training, inference optimization, and large-scale distributed compute for Tenstorrent's custom AI accelerators, using Python, PyTorch, and techniques like speculative decoding and distributed training.
Research scientist at an MIT-born, venture-backed startup building an AI copilot for design and manufacturing. You'll develop geometry processing and simulation algorithms, build services for 2D/3D CAD/CAE/CAM data, create datasets and rendering modules, and optimize ML post-training workflows. Core stack is Python and machine learning.
Principal Software Development Engineer at Zscaler (hybrid, Bangalore) on the ZPA team, acting as a hands-on technical leader for large-scale Generative AI platforms and agent systems in production. Day to day involves building MCP servers, deploying containerized AI workloads with Docker/Kubernetes, and defining engineering standards for enterprise AI.
ML Systems Engineer at Final, an HFT/trading-algorithms firm in Ramat Hasharon, building and optimizing proprietary deep learning systems for live trading. Day-to-day involves running DL models on large GPU clusters and adapting them for production serving, using PyTorch, Python/C/C++, and custom CUDA/Triton kernels.
Leads a small team optimizing machine learning model performance for autonomous driving software, focusing on C++/Python code and ML compiler tools.
Senior/Staff ML Engineer at Waymo building ultra-realistic 3D/4D world models and generative systems for autonomous vehicle simulation using advanced ML techniques like diffusion models and VLMs.
A senior machine learning engineer at Waymo optimizes generative models and simulations for autonomous driving, focusing on performance, efficiency, and scalability using advanced ML techniques and hardware accelerators.
We couldn't check your fit for this role — add a CV to your profile to see it next time.