Senior AI Storage Infrastructure Engineer
Summary
Architect and maintain high-performance storage infrastructure for AI training/inference workloads, building CSI drivers, GPUDirect Storage integrations, NVMe caching layers, and RDMA networking on Kubernetes.
You will architect the storage data-delivery fabric for intensive AI training and inference workloads. You will develop CSI drivers, integrate GPUDirect Storage, manage NVMe caching, tune I/O performance, optimize RDMA connectivity, implement storage monitoring and quotas, and enforce multi-tenant resource isolation.
Responsibilities
- Design and maintain CSI drivers for high-performance parallel file systems
- Implement GPUDirect Storage integrations
- Develop local NVMe caching strategies for models and datasets
- Optimize IOPS throughput and latency across the storage stack
- Integrate storage with RDMA InfiniBand and RoCE networks
- Implement storage performance monitoring and alerting
- Define storage policies quotas and Kubernetes multi-tenancy isolation
- Mentor engineers and lead architectural design reviews
Requirements
- Bachelor’s or Master’s degree in Computer Science Electrical Engineering or a related field
- 5+ years of experience with distributed storage and high-performance file systems
- Understanding of POSIX compliance and file I/O semantics
- Deep expertise in Kubernetes CSI including volume plugins and storage operators
- Experience with Linux block and file I/O and kernel performance tuning
- Familiarity with RDMA InfiniBand and RoCE
- Experience operating debugging and scaling production or HPC storage environments
- Experience with Terraform Ansible and CI/CD pipelines
- Strong technical communication skills