Point your AI agent at freehire and let it find you a job.

Get the CLI →

World Wide Technology

NewBe an early applicant

Domain Architect AI Compute

Posted Updated
Discussion

About the role

The Domain Architect - AI Compute acts as the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across diverse client environments, bridging the gap between architectural design and hands-on execution. Operating with a 60/40 split between delivering complex AI infrastructure (60%) and providing Pre-Sales Subject Matter Expertise (40%), you will lead the physical provisioning of NVIDIA SuperPOD, NVIDIA BasePOD, and Cisco AI Factory environments, ensuring clients receive "Day 2" ready AI factories, while assisting the sales team in defining the scope and cost of future deployments.

Key responsibilities

  • Lead the physical provisioning of clusters: NVIDIA NVL72, DGX SuperPOD, BasePOD, HGX, MGX, Cisco AI Factory

  • Utilise NVIDIA Base Command Manager (BCM) for diskless booting, firmware management, and OS hardening

  • Establish monitoring with NVIDIA Mission Control

  • Execute automated "Zero Touch Provisioning" (ZTP) workflows to transform bare-metal hardware into production-ready nodes

  • Define and enforce "Fair Share" policies, fractional GPU quotas using Multi-Instance GPU (MIG), and pre-emption logic for multi-tenant environments

  • Implement and configure advanced schedulers: Slurm (bare metal), NVIDIA Run:AI (Kubernetes), Kueue (Kubernetes), Volcano (Kubernetes)

  • Deploy and configure management planes like Rafay or Armada to enable multi-cluster management and observability

  • Implement high-fidelity telemetry using DCGM (Data Centre GPU Manager) to monitor GPU health, thermal throttling, and XID error rates

  • Conduct validation testing using NCCL-tests, HPL, and HPCG to verify cluster performance

  • Assist the sales team by validating customer technical requirements and producing accurate Labour Estimates (LOE) for Statements of Work (SOWs)

About you

  • Deep architectural understanding of NVIDIA GPU platforms (Hopper, Grace-Hopper, Blackwell, Grace-Blackwell)

  • Mastery of NVL72 rack-scale integration and NVSwitch fabrics

  • Expertise in the associated software stack (CUDA, cuDNN, NCCL)

  • Expert-level knowledge of Linux distributions (Ubuntu, RHEL) optimised for HPC/AI

  • Deep experience with kernel tuning, driver management, and system hardening

  • Proficiency in Python and Ansible for hardware configuration management and automation

  • Experience working within a System Integrator (SI) or Managed Service Provider (MSP) environment (desirable)

  • Hands-on experience with NVIDIA Base Command Manager (BCM) and/or NVIDIA Mission Control (desirable)

  • Solid understanding of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2) and how they interface with host PCIe/NVLink topologies (desirable)

  • Experience with Kubernetes/Red Hat OpenShift installation and administration (desirable)


Skills

See also

Architecture jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available