Point your AI agent at freehire and let it find you a job. A CLI and an MCP server over the whole job API — no browser.
Навыки: Docker, Cuda, CI/CD, Kubernetes, LLM. Специализации: MLOps-инженер. Хотите стать частью увлекательного процесса цифровой трансформации? Блок IT в СОГАЗ активно развивается и меняет подход к созданию продуктов.…
Builds and prototypes AI integrations for Nebius’s cloud platform, shaping reference architectures and product requirements while collaborating with partner engineering teams.
Build and optimize GPU infrastructure for AI workloads, profiling performance across hardware and frameworks to guide platform decisions and hardware development.
Senior ML Engineer builds and optimizes low-level GPU inference kernels and runtime components for a large-scale AI cloud platform, focusing on performance tuning and hardware integration.
Senior ML Engineer at Nebius builds and optimizes high-performance inference and fine-tuning platforms for large language models across tens of thousands of GPUs, focusing on throughput, latency, and cost-per-token.
The Senior ML Solutions Architect will help customers design and implement optimized LLM inference and fine-tuning workflows using the Nebius Token Factory platform. This role involves architecting scalable AI applications, providing technical expertise in RAG and prompt engineering, and collaborating with internal engineering teams to improve the platform.
Senior ML Solutions Architect designs and optimizes serverless LLM inference and fine-tuning workflows for customers using Nebius Token Factory, focusing on production deployment and scalability.
Nebius is seeking a Senior Sales Engineer to act as a technical partner for customers using their AI inference platform. The role involves leading technical discovery, validating PoC architectures, and ensuring that customer AI workloads are scalable and economically viable on their GPU-backed infrastructure.
Senior Sales Engineer at Nebius designs and validates scalable AI inference architectures for enterprise customers, bridging technical feasibility with commercial growth in a global AI cloud platform.
Build and maintain the reliability, observability, and performance of Nebius’s AI inference platform, ensuring seamless operation under extreme load while optimizing GPU workloads and Kubernetes clusters.
Builds and optimizes low-level GPU kernels and runtime components for an AI inference platform, integrating new hardware and collaborating with ML teams to improve performance at scale.
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through…
Early-career ML Solutions Architect builds and tests LLM-based applications, benchmarks models, and optimizes inference on a serverless AI platform, learning scalable AI deployment with mentorship from senior architects.
Builds and optimizes a cloud platform for serving large-scale AI models efficiently, integrating open-source frameworks like vLLM and TRT-LLM with advanced systems techniques.
Principal ML Solutions Architect designs and optimizes serverless LLM inference and fine-tuning workflows for enterprise customers, using frameworks like vLLM and TensorRT-LLM on Nebius's AI cloud platform.
Lead a red team to simulate advanced cyberattacks against a cutting-edge AI cloud platform, uncovering vulnerabilities in GPU clusters, inference stacks, and multi-tenant infrastructure to drive security improvements.
Senior Applied Scientist at Nebius designs and optimizes efficient LLM/VLM inference methods, turning research into production systems using PyTorch, CUDA, and Triton.
Senior Machine Learning Engineer at Nebius in Palo Alto optimizing LLM inference systems for performance, cost, and reliability using frameworks like vLLM and PyTorch.
Nebius is seeking a Senior Software Engineer to build and scale their GPU-native Serverless AI platform. You will design core components like schedulers and runtimes, optimize for low-latency performance, and lead technical architecture for AI infrastructure.
We couldn't check your fit for this role — add a CV to your profile to see it next time.