Domain Architect (AI Storage)
About the role
The Domain Architect - AI Storage acts as the primary technical authority for the physical and logical lifecycle of high-performance data platforms across diverse client environments, bridging the gap between architectural design and hands-on execution. You will define the "Gold Standard" for storage infrastructure, moving beyond single-array management to architect repeatable, scalable, and automated data fabrics. You will serve as the technical lead for NVIDIA Cloud Provider (NCP) and private enterprise AI cloud deployments, owning the "Storage" in the critical "Compute-Network-Storage" triad. You will operate with a 60/40 split between delivering complex AI infrastructure (60%) and providing Pre-Sales Subject Matter Expertise (40%).
Key responsibilities
Lead the installation and configuration of high-performance storage clusters using technologies such as WEKA, DDN (Lustre), VAST Data, or Pure Storage
Optimise storage client configurations on compute nodes, managing kernel modules and mounting parameters to ensure stability at scale
Implement GPUDirect Storage (GDS) technologies to bypass the CPU and enable direct data paths between NVMe drives and GPU memory
Deploy and configure Container Storage Interface (CSI) drivers for Kubernetes/Red Hat OpenShift, ensuring persistent storage is dynamically provisioned for AI workloads
Design storage classes that differentiate between "Scratch" (High Performance) and "Home/Project" (General Purpose) tiers
Execute synthetic benchmark suites (IOR, FIO, mdtest) to validate throughput and metadata performance against agreed SLAs
Troubleshoot I/O issues, analysing client-side logs and fabric counters
Implement automated data tiering strategies to move datasets between Hot (NVMe), Warm (QLC Flash), and Cold (Object/S3/Tape) tiers
Lead the architectural sizing for storage opportunities and calculate required performance for specific workloads
Own the Storage Bill of Materials (BoM), ensuring the correct ratio of Storage Servers to Compute Nodes/GPUs
About you
Expert-level knowledge of Parallel File Systems (WEKA, Lustre, BeeGFS, GPFS)
Deep understanding of Object Storage (S3 protocols) for model checkpointing and archiving
Mastery of storage protocols: NVMe-over-Fabrics (NVMe-oF), NFS over RDMA, and NVIDIA GPUDirect Storage (GDS)
Deep understanding of the Linux I/O stack, including Block Device drivers, file system tuning (xfs, ext4), and client-side caching mechanisms
Proficiency in one of Python, Ansible, or Terraform for automating storage cluster deployment and client configuration management
Experience working within a System Integrator (SI), Storage Vendor (e.g., NetApp, Dell, Pure), or MSP environment
Hands-on experience integrating storage with bare metal, and Kubernetes/Red Hat OpenShift
Understanding of InfiniBand and RoCEv2 fabrics from a storage perspective (Congestion Control, Quality of Service)