Tech jobs
Job listings
Technical Account Manager - Token Factory
About Nebius: Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through…
Technical Product Manager - Soperator
Own the product direction for Soperator, Nebius's Slurm-on-Kubernetes control plane for GPU clusters, shaping how ML engineers run and scale distributed AI workloads using cloud infrastructure and orchestration tools.
Technical Project Manager - IAAS
Technical Project Manager leading cross-team initiatives to build and deliver IaaS cloud infrastructure for AI workloads, coordinating compute, networking, storage, Kubernetes, and APIs.
Technical Project Manager - Security
Oversee security-focused cloud infrastructure projects, coordinating cross-team efforts to implement and deliver security programs for an AI cloud platform.
Staff / Senior Software Engineer (Agentic Search) - Crawler
Builds and operates web-scale crawling infrastructure for an AI-native search platform, processing billions of URLs to feed real-time data to agentic AI systems using Go/C++ and distributed systems.
Senior Site Reliability Engineer (In-Office Required)
Senior Site Reliability Engineer responsible for designing, scaling, and maintaining Kubernetes-based infrastructure and CI/CD pipelines for an AI-native company powering agentic web interactions.
Senior Software Developer: Models Team (Token Factory)
Builds and optimizes a cloud platform for serving large-scale AI models efficiently, integrating open-source frameworks like vLLM and TRT-LLM with advanced systems techniques.
Technical Product Manager – Storage
Owns the storage product vision and roadmap for a cloud AI platform, defining technical requirements for block, file, object, and HPC storage while balancing performance, durability, and cost.
Forward Deployed Engineer - Physical AI Cloud Platform
Senior engineer embedded with customers to design, build, and deploy cloud infrastructure that runs physical AI workloads like simulation, training, and inference on Nebius's AI cloud platform.
Site Reliability Engineer in Hardware Infrastructure
Senior Site Reliability Engineer at Nebius, an AI cloud platform company, automating and scaling data center hardware infrastructure using Linux, Python, and Bash to ensure fault-tolerant, high-performance systems.
Senior Machine Learning Engineer, Model Training and Reinforcement Learning
Senior ML Systems Engineer builds and maintains distributed AI training infrastructure for large-scale model training and reinforcement learning experiments using PyTorch, Megatron-LM, and DeepSpeed in Palo Alto.
Senior Machine Learning Engineer, LLM Inference Optimization
Senior Machine Learning Engineer at Nebius in Palo Alto optimizing LLM inference systems for performance, cost, and reliability using frameworks like vLLM and PyTorch.
Senior Technical Project Manager - VPC
Senior Technical Project Manager to lead cloud networking initiatives for Nebius AI Cloud, focusing on VPC development, scalability, and cross-team coordination in a fast-growing AI infrastructure company.
Senior Software Engineer (Agentic Search) - Web Access Engineer
Builds and maintains browser infrastructure and tooling for a novel search engine designed for AI agents, focusing on reliable web access at internet scale using modern browser technologies and automation frameworks.
Senior Software Engineer (Agentic Search) – Web Rendering Engineer
Builds and optimizes large-scale browser rendering infrastructure to extract structured data from modern web apps for AI agents, using Chromium internals, headless browsers, and frontend frameworks.
Senior ML Engineer (AI Research/ Portability)
Build and evaluate AI agent systems that remain portable across models, providers, and deployment environments, focusing on routing, memory, interoperability, and evaluation.
DevOps Engineer (Agentic Search)
Build and run the Kubernetes-based infrastructure that powers Tavily’s agentic search API, including CI/CD, observability, and multi-region deployments handling billions of daily events.
Principal DevSecOps Engineer - AI Infrastructure
Builds and maintains the core AI infrastructure platform, ensuring scalability, reliability, and security for a high-growth AI company’s distributed systems using IaC, GitOps, and coding in Go/Python/Rust.
Principal Software Engineer
Leads technical design and presales for Java-based distributed systems, translating business needs into scalable microservices architectures, cloud solutions, and enterprise-grade implementations while mentoring engineers and guiding project kickoffs.