Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Build and lead large-scale AI systems for federal programs, designing Python-based LLM solutions, RAG pipelines, and cloud-native architectures while ensuring compliance and security in regulated environments.
Maintains and optimizes a hybrid HPC/AI Linux cluster with GPUs, scheduling tools (SLURM/Kubernetes), and MLOps pipelines to support large-scale model training and inference for researchers.
Build and deploy large-scale machine learning models and AI agents to optimize Micron’s semiconductor manufacturing workflows using distributed training and GPU optimization techniques.
Design and optimize ML pipelines, LLM serving, and GPU architectures for generative AI solutions using frameworks like PyTorch, Hugging Face, and vLLM.
Build and improve Avito’s in-house large language models by designing experiments, training pipelines, and optimizing inference for production use.
Build and scale AI/ML infrastructure for autonomous-driving simulations using large foundation models, focusing on distributed training and ML accelerators.
Builds and optimizes large-scale AI infrastructure, including Kubernetes clusters, RDMA networking, and GPU orchestration to improve efficiency and scalability of AI training/inference systems.
Build and optimize multi-agent LLM systems, RAG pipelines, and large-scale NLP models to transform enterprise processes with rapid prototyping and production deployment.
R&D-focused NLP role building multi-agent LLM systems, RAG pipelines, and optimizing inference for large-scale text processing in a fintech environment.
Разрабатываете NLP/мультимодальные пайплайны, RAG-системы и ИИ-агентов для бизнес-задач, интегрируете их в высоконагруженные сервисы банка.
Build and deploy large language models for financial workflows, optimizing training and serving to improve agent efficiency and customer experience in banking systems.
Build and optimize large language models for financial services, focusing on LLM-based methods, training pipelines, and production deployment to improve customer workflows and agent efficiency.
Build and optimize distributed GPU training infrastructure for large neural networks, focusing on PyTorch, Megatron-LM, and DeepSpeed to maximize throughput and stability at scale.
Build and improve an AI shopping assistant that surfaces relevant listings, compares options, and suggests fair prices for millions of users using modern ML and LLM techniques.
Build and ship state-of-the-art voice conversion and speech models end-to-end, from data curation to production inference, using large-scale diffusion/flow-matching transformers and PyTorch.
Build and scale the distributed training infrastructure for a cutting-edge video-generation AI model, optimizing multi-GPU clusters and data pipelines for stability and performance.
Design and build the enterprise AI platform’s core architecture, reusable patterns, and agentic AI systems for New York Life’s insurance and financial services.
Design and optimize large-scale AI training and inference clusters for enterprise clients, benchmarking performance and creating reference architectures for NVIDIA GPU-based systems.
Principal Software Engineer, E2E Performance and Goodput — CSP Engagements Location: US, CA, Santa Clara, US, TX, Austin, US, OR, Remote, US, CA, Remote Time Type: Full time Job Description We're looking for a…
Lead end-to-end design and delivery of large-scale AI systems for federal programs, building Python-based LLM and generative AI solutions with cloud-native architectures and MLOps pipelines.
We couldn't check your fit for this role — add a CV to your profile to see it next time.