Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and run a production-grade ML platform: design AI system architectures, deploy and optimize LLM inference servers, and maintain MLOps pipelines with GPU scheduling and monitoring.
Build, deploy, and optimize production AI systems including LLMs and generative AI models, focusing on inference pipelines, performance, and cost efficiency.
Lead architecture and deployment of production AI systems, optimizing inference pipelines and mentoring engineers in a fast-growing company.
Lead a team to build and deploy large language model applications, fine-tune models, and design RAG systems for mission-critical government use cases.
Optimize distributed AI training for 100B+ parameter models across 100k+ GPUs, writing CUDA/Triton kernels and co-designing hardware-efficient training recipes.
Build production-grade AI architecture blueprints and open-source Quickstarts with Python, PyTorch, and Kubernetes, focusing on enterprise deployment, security, and regulated environments.
Build and maintain a secure AI platform for law-enforcement investigations, integrating LLMs, RAG, and agentic workflows with backend services and frontend interfaces.
DevOps engineer building and operating Kubernetes-based AI chat platforms, CI/CD pipelines, and Kafka data pipelines while also developing NestJS microservices for chat, rewards, and analytics.
Build and optimize distributed systems that run large language models efficiently across Intel hardware, focusing on inference performance, parallelism, and communication.
Senior engineer building and optimizing open-source AI frameworks (Megatron Core, NeMo) for large language and multimodal model training, fine-tuning, and deployment on NVIDIA GPUs.
Mid-level AI developer building and integrating locally hosted LLMs and ML models into classified government systems, focusing on secure, air-gapped environments and production-grade AI features.
Mid-level AI developer building and integrating locally hosted LLMs and ML models into classified government systems, focusing on secure, air-gapped environments and production-grade delivery.
Lead enterprise GenAI strategy, architect scalable AI systems, and advise C-suite on tech stacks, trade-offs, and ROI for large client engagements.
Build and optimize production-grade AI systems using RAG pipelines, vector stores, and cloud-native tools to deliver secure, scalable solutions for federal environments.
Обязанности: Разрабатывать системы анализа здоровья: наше ключевое направление AI, который помогает пользователю понять своё состояние и вовремя дойти до врача; Внедрять AI-фичи в продукт и в компанию: интеграция LLM…
Build and deploy quantized large language models for in-vehicle AI, focusing on PTQ, QAT, and low-bit inference to ensure numerical consistency and performance on XPENG’s Turing AI chip.
Build and deploy quantized large language models for in-vehicle AI, focusing on PTQ, QAT, and low-bit inference to optimize performance on XPENG’s Turing AI chip.
Designs and validates AI accelerator systems (e.g., Gaudi, GPUs) for large-scale ML workloads, debugging hardware, firmware, and software layers while leading cross-functional teams.
Develops and optimizes distributed AI infrastructure and software for LLMs, vision AI, and robotics using Intel Xeon/GPU hardware.
Develop and optimize PyTorch-based AI frameworks for Intel hardware, focusing on distributed algorithms, kernel fusion, and GPU enablement to boost model performance.
We couldn't check your fit for this role — add a CV to your profile to see it next time.