Senior AI Storage Infrastructure Engineer

Summary

Architect and maintain high-performance storage infrastructure for AI training/inference workloads, building CSI drivers, GPUDirect Storage integrations, NVMe caching layers, and RDMA networking on Kubernetes.

You will architect the storage data-delivery fabric for intensive AI training and inference workloads. You will develop CSI drivers, integrate GPUDirect Storage, manage NVMe caching, tune I/O performance, optimize RDMA connectivity, implement storage monitoring and quotas, and enforce multi-tenant resource isolation.

Responsibilities

  • Design and maintain CSI drivers for high-performance parallel file systems
  • Implement GPUDirect Storage integrations
  • Develop local NVMe caching strategies for models and datasets
  • Optimize IOPS throughput and latency across the storage stack
  • Integrate storage with RDMA InfiniBand and RoCE networks
  • Implement storage performance monitoring and alerting
  • Define storage policies quotas and Kubernetes multi-tenancy isolation
  • Mentor engineers and lead architectural design reviews

Requirements

  • Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
  • 5+ years of experience with distributed storage and high-performance file systems
  • Understanding of POSIX compliance and file I/O semantics
  • Deep expertise in Kubernetes CSI including volume plugins and storage operators
  • Experience with Linux block and file I/O and kernel performance tuning
  • Familiarity with RDMA InfiniBand and RoCE
  • Experience operating debugging and scaling production or HPC storage environments
  • Experience with Terraform Ansible and CI/CD pipelines
  • Strong technical communication skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available