GPU Infrastructure Engineer
Summary
Build and operate a distributed multi-node GPU platform for large-scale AI model training and inference, ensuring performance and reliability using Kubernetes, Slurm, and NVIDIA stack technologies.
We’re looking for an experienced Infrastructure Engineer to work on a project of our Customers. They are building a multi-node GPU platform for large-scale model training and inference using managed GPU providers. In this role, you’ll have to own the platform performance, reliability, and the technical relationship with the Client’s providers.
Our Customers provide SaaS solutions that help companies to optimize their business. These solutions include business planning to automate and optimize business, delivery, and workflow solutions. The platform leverages industry-leading Artificial Intelligence (AI) and Machine Learning (ML) for better prediction and prevention of disruptions across business.
Candidate’s location – Europe.
Responsibilities:
Design and operate distributed GPU training infrastructure
Validate cluster topology, RDMA/InfiniBand, and NCCL performance
Standardize Kubernetes or Slurm scheduling, GPU images, and software versions
Diagnose issues across training workloads, networking, storage, and GPU hosts
Build monitoring, benchmarks, runbooks, and reliability standards
Work with GPU providers to resolve incidents and define technical requirements
Support SFT, DPO, RL, and large-scale inference workloads
Requirements:
Location in Europe
Production experience with multi-node GPU training infrastructure
Strong Linux, containers, CUDA, and NVIDIA-GPU-stack knowledge
Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
Deep experience with Kubernetes or Slurm
Infrastructure automation and observability experience
Strong skills in incident leadership and provider-facing communication
English level – Upper-Intermediate or higher
Will be a plus:
Experience in an AI lab, HPC environment, or specialist GPU cloud
Knowledge of distributed-training frameworks such as PyTorch, Megatron, or DeepSpeed
Experience with parallel storage, checkpoint optimization, and multi-provider platforms