Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Builds and maintains developer-facing tools (APIs, SDKs, dashboards) on top of AI inference/training infrastructure like SGLang and Miles, collaborating with product and research teams to create intuitive interfaces.
Build and optimize production-grade inference services for Apple’s large language, vision, and speech models, ensuring low-latency performance across products like Siri and Photos.
Build and optimize distributed training systems for large neural networks (LLMs, diffusion, SSMs) across GPU clusters, focusing on throughput, stability, and fault tolerance using PyTorch, Megatron-LM, and DeepSpeed.
Build and optimize low-latency ML inference pipelines and LLM tools to automate trading workflows, using PyTorch, TensorRT, and cloud GPUs.
Build and optimize large language models for insurance workflows, focusing on post-training, evaluation, and inference performance to improve underwriting and claims processing.
Principal AI/ML Architect designs and advises on production ML systems, MLOps/LLMOps pipelines, and GenAI architectures on AWS for enterprise clients, translating technical depth into business value.
Optimize AI training and inference workloads for speed, cost, and efficiency across the full stack, from GPU kernels to distributed systems, using Python, C++, and profiling tools.
Build and deploy agentic AI systems for financial data, focusing on LLM orchestration, retrieval, and scalable workflows to power research and insights.
Build and deploy agentic AI systems for financial data, focusing on LLM orchestration, retrieval, and scalable workflows to power generative AI applications in finance.
Build and optimize distributed AI training and inference pipelines for gaming using PyTorch, DeepSpeed, and Kubernetes to support large language models and reinforcement learning workloads.
Lead end-to-end training of large language models using domain-adaptive pretraining, fine-tuning, and reinforcement learning, while building robust data and evaluation pipelines for intelligent agent systems.
Build and fine-tune large language models for Avito’s products, optimizing training pipelines and inference speed for production-scale NLP systems.
Builds low-level system software to optimize distributed AI training and inference across thousands of GPUs using C++, Python, and CUDA.
ML engineer builds and runs online reinforcement-learning pipelines to improve GigaChat’s post-training, designing experiments, reward signals, and distributed training workflows.
Build and own the shared AI platform that trains and serves Adobe’s generative AI models at global scale, focusing on GPU fleet utilization, low-latency inference, and end-to-end model deployment pipelines.
Lead a team building and optimizing scalable ML training platforms on AWS/Kubernetes, focusing on GPU workloads, performance tuning, and Gen AI/LLM pipelines while enforcing enterprise governance and security standards.
Principal AI/ML Research Engineer leads applied research in generative AI, deep learning, and Transformers to build Payment Foundation Models for commerce and fintech challenges, driving novel architectures from prototype to production.
About us We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training,…
Build and operate the distributed orchestration engine for Amazon SageMaker AI’s Model Factory, running large-scale LLM training and customization workflows across thousands of GPUs and Trainium devices.
We couldn't check your fit for this role — add a CV to your profile to see it next time.