Tech jobs
Job listings
Machine Learning - Senior Software Engineer - Python, DevOps
Senior ML engineer who builds, deploys, and optimizes production ML pipelines, focusing on Python-based inference services, TensorRT acceleration, and DevOps automation (CI/CD, IaC).
Senior Site Reliability Engineer (In-Office Required)
Owns all infrastructure for an AI-native API company—managing Kubernetes clusters, IaC (Terraform), CI/CD/GitOps pipelines, real-time data pipelines processing billions of events, and observability stacks in a fully in-office NYC role.
AI/ML Specialist Solutions Architect
Designs and advises on scalable AI cloud solutions for enterprise customers, focusing on distributed training and inference pipelines using PyTorch, JAX and Kubernetes.
Senior/Staff Infrastructure Security Engineer
Senior engineer secures Nebius’ AI cloud platform by hardening Kubernetes, Linux, networks, and cloud infrastructure while automating security controls and responding to threats.
Key Customers Solutions Architect
Designs and deploys large-scale AI/ML workloads on GPU cloud infrastructure, advising key customers and optimizing performance while collaborating with sales and product teams.
Key Customers Solutions Architect EMEA
Solutions Architect at a cloud platform company, advising enterprise AI customers on deploying and scaling GPU workloads for ML training and inference using Nebius’s AI cloud services.
Partner Solutions Architect
Partner Solutions Architect designs and implements integrations between Nebius' AI cloud platform and strategic partners, enabling joint customers to deploy GPU-accelerated AI/ML solutions at scale.
Senior Site Reliability Engineer — Token Factory (Inference Platform)
Build and maintain the reliability, observability, and performance of Nebius’s AI inference platform, ensuring seamless operation under extreme load while optimizing GPU workloads and Kubernetes clusters.
Senior Technical Project Manager (Hardware Automation)
Leads cross-team automation projects to streamline hardware operations in a cloud AI infrastructure company, coordinating engineering, operations, and infrastructure teams to reduce manual effort at scale.
Site Reliability Engineer (SRE) AI Infrastructure (Early Career)
Early-career Site Reliability Engineer supporting AI cloud infrastructure operations, deploying changes, automating tasks, and learning Kubernetes, networking, and cloud systems.
Senior Site Reliability Engineer (In-Office Required)
Senior Site Reliability Engineer responsible for designing, scaling, and maintaining Kubernetes-based infrastructure and CI/CD pipelines for an AI-native company powering agentic web interactions.
Infrastructure Site Reliability Engineer
Build and run the core network infrastructure for a cloud platform powering AI workloads, ensuring high availability, observability, and safe automation at scale.
Software Engineer in Infrastructure
Designs and builds network automation and observability tooling for a global AI cloud platform, using Go/Python to make large-scale data center networks safe, scalable, and reliable.
Principal ML Solutions Architect - Token Factory
Principal ML Solutions Architect designs and optimizes serverless LLM inference and fine-tuning workflows for enterprise customers, using frameworks like vLLM and TensorRT-LLM on Nebius's AI cloud platform.
Forward Deployed Engineer - Physical AI Cloud Platform
Senior engineer embedded with customers to design, build, and deploy cloud infrastructure that runs physical AI workloads like simulation, training, and inference on Nebius's AI cloud platform.
Cloud Solution Architect - Educational Content Author, Nebius Academy
Designs and creates technical tutorials, sample code, and reference architectures to teach developers how to use Nebius' AI cloud infrastructure, including VMs, GPU clusters, and Kubernetes.
Partner Solutions Architect
Designs and implements AI/ML cloud integrations with strategic partners, enabling joint customers to deploy GPU-accelerated solutions efficiently using modern frameworks and cloud infrastructure.
Site Reliability Engineer in Network Infrastructure
Build and run the core network infrastructure for a cloud platform powering AI workloads, defining reliability targets and automating operations to scale safely.
Software Engineer in Network Infrastructure
Builds and maintains network automation and observability tooling for a cloud platform powering AI workloads, using Go/Python to make large-scale network operations safe and scalable.