Point your AI agent at freehire and let it find you a job.

Get the CLI →

Launch & Scale

NewBe an early applicant

Senior Infrastructure Engineer (GPU Platform)

Discussion

Summary

Build and maintain multi-node GPU training infrastructure using Linux, containers, CUDA, and Kubernetes/Slurm, focusing on networking, storage, and observability for AI workloads.

On behalf of our client, we are looking for a Senior Infrastructure Engineer (GPU Platform)

Requirements:

  • Production experience with multi-node GPU training infrastructure
  • Strong Linux, containers, CUDA, and NVIDIA GPU stack knowledge
  • Hands-on experience with NCCL and InfiniBand or RoCE/RDMA troubleshooting
  • Deep experience with Kubernetes or Slurm
  • Experience with infrastructure automation and observability
  • Experience diagnosing issues across training workloads, networking, storage, and GPU hosts
  • Strong incident leadership and provider-facing communication skills
  • English – Upper-Intermediate or higher

Would be a plus:

  • Experience in an AI lab, HPC environment, or specialist GPU cloud
  • PyTorch, Megatron, DeepSpeed, or other distributed-training frameworks
  • Experience with parallel storage and checkpoint optimization
  • Experience working with multi-provider GPU platforms

Company offers:

  • Long-term employment with possibilities for professional growth
  • Fully remote work
  • Reasonably flexible schedule
  • 15 days of paid vacation
  • Regular performance reviews

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available