Domain Architect AI Compute
About the role
The Domain Architect - AI Compute acts as the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across diverse client environments, bridging the gap between architectural design and hands-on execution. Operating with a 60/40 split between delivering complex AI infrastructure (60%) and providing Pre-Sales Subject Matter Expertise (40%), you will lead the physical provisioning of NVIDIA SuperPOD, NVIDIA BasePOD, and Cisco AI Factory environments, ensuring clients receive "Day 2" ready AI factories, while assisting the sales team in defining the scope and cost of future deployments.
Key responsibilities
Lead the physical provisioning of clusters: NVIDIA NVL72, DGX SuperPOD, BasePOD, HGX, MGX, Cisco AI Factory
Utilise NVIDIA Base Command Manager (BCM) for diskless booting, firmware management, and OS hardening
Establish monitoring with NVIDIA Mission Control
Execute automated "Zero Touch Provisioning" (ZTP) workflows to transform bare-metal hardware into production-ready nodes
Define and enforce "Fair Share" policies, fractional GPU quotas using Multi-Instance GPU (MIG), and pre-emption logic for multi-tenant environments
Implement and configure advanced schedulers: Slurm (bare metal), NVIDIA Run:AI (Kubernetes), Kueue (Kubernetes), Volcano (Kubernetes)
Deploy and configure management planes like Rafay or Armada to enable multi-cluster management and observability
Implement high-fidelity telemetry using DCGM (Data Centre GPU Manager) to monitor GPU health, thermal throttling, and XID error rates
Conduct validation testing using NCCL-tests, HPL, and HPCG to verify cluster performance
Assist the sales team by validating customer technical requirements and producing accurate Labour Estimates (LOE) for Statements of Work (SOWs)
About you
Deep architectural understanding of NVIDIA GPU platforms (Hopper, Grace-Hopper, Blackwell, Grace-Blackwell)
Mastery of NVL72 rack-scale integration and NVSwitch fabrics
Expertise in the associated software stack (CUDA, cuDNN, NCCL)
Expert-level knowledge of Linux distributions (Ubuntu, RHEL) optimised for HPC/AI
Deep experience with kernel tuning, driver management, and system hardening
Proficiency in Python and Ansible for hardware configuration management and automation
Experience working within a System Integrator (SI) or Managed Service Provider (MSP) environment (desirable)
Hands-on experience with NVIDIA Base Command Manager (BCM) and/or NVIDIA Mission Control (desirable)
Solid understanding of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2) and how they interface with host PCIe/NVLink topologies (desirable)
Experience with Kubernetes/Red Hat OpenShift installation and administration (desirable)