Enterprise Architect AI
Summary
World Wide Technology is hiring a hands-on Enterprise Architect for AI in London (remote with up to 30% client-site travel) to own end-to-end delivery of client AI engagements — from GPU/data-center and high-performance networking design through training/inference pipelines, MLOps/LLMOps tooling, and AI governance frameworks.
Salary: £62,000 - 102,000 per year
Requirements:- 10+ years in enterprise architecture, infrastructure engineering, or platform engineering roles.
- 5+ years focused specifically on AI/ML systems design and delivery, including at least 2 years working with generative AI/LLM workloads.
- Demonstrated track record leading technical delivery, not just advisory, on enterprise-scale AI or HPC infrastructure programmes.
- Experience with GPU/accelerator architectures, including NVIDIA or AMD, and multi-node scale-out design.
- Experience with accelerator interconnects such as NVLink and NVSwitch.
- Experience with high-performance networking, including InfiniBand, RoCEv2 fabric design, 400G/800G Ethernet, and rail-optimized topologies for AI clusters.
- Experience with data center facilities considerations such as power density, liquid cooling, and rack-level design for AI compute.
- Experience with parallel and high-throughput file systems sized for training and checkpointing workloads.
- Hands-on experience with Kubernetes, Docker, Slurm, Run:ai, or equivalent GPU scheduling platforms.
- Working-level experience with PyTorch and TensorFlow.
- Experience with distributed training frameworks such as Horovod, DeepSpeed, or Megatron-LM.
- Experience with inference and serving platforms such as NVIDIA Triton, vLLM, or TensorRT-LLM.
- Experience with MLOps/LLMOps tools including Kubeflow, MLflow, and at least one hyperscaler ML platform.
- Experience with generative AI techniques including LLM fine-tuning, RAG architecture design, vector databases, and agentic frameworks.
- Experience with infrastructure-as-code tools such as Terraform and Ansible, and GitOps with ArgoCD.
- Experience with pipeline orchestration tools such as Kubeflow Pipelines, Apache Airflow, or Argo Workflows.
- Experience with GPU job scheduling and Kubernetes-native GPU scheduling, including device plugins and MIG partitioning.
- Experience with CI/CD/CT for ML, including automated model testing, validation gates, and promotion pipelines.
- Experience with infrastructure and GPU observability tools such as NVIDIA DCGM, Prometheus/Grafana, and related telemetry stacks.
- Experience with model and LLM observability, including performance monitoring, drift detection, and tools such as Arize, WhyLabs, or Langfuse.
- Experience with centralized logging and distributed tracing across data, training, and inference pipelines.
- Experience integrating AI platforms with enterprise systems using REST/gRPC APIs and message/event streaming platforms.
- Experience designing connectivity between AI/GPU infrastructure and enterprise network environments.
- Experience integrating on-premises AI platforms with cloud AI services via dedicated interconnects and hybrid/multi-cloud connectivity patterns.
- Experience architecting connections into colocation interconnects, managed GPU-as-a-service offerings, and external inference endpoints.
- Working knowledge of API gateways, service mesh, and mutual TLS for secure exposure of AI services.
- Working knowledge of model risk management frameworks and responsible AI principles.
- Familiarity with data privacy regulation as applied to AI training and inference data.
- Working knowledge of emerging AI-specific regulation and standards such as the EU AI Act, NIST AI RMF, and ISO/IEC 42001.
- Experience establishing model documentation, audit trails, and approval-gate processes for production AI systems.
- AI-specific security fundamentals, including model security, prompt-injection defenses, and supply chain security for open-source/open-weight models.
- Solutions-architect level expertise in at least one hyperscaler, including native AI/ML services.
- Ability to design for hybrid on-premises/cloud AI deployments, including data residency and sovereignty constraints.
- Proven experience hosting and chairing formal Architecture Review Board and Technical Design Authority forums.
- Ability to define and operate governance gates across the engagement lifecycle.
- Experience producing and maintaining architecture decision records, design standards, and reference architectures.
- Excellent executive communication and presentation skills.
- Proven ability to author low-level design and high-level design documentation to a professional services standard.
- Experience running technical design workshops with senior client stakeholders and multi-vendor delivery teams.
- Bachelors degree in Computer Science, Computer Engineering, or a related technical field, or equivalent demonstrable experience.
- Advanced degree in CS, AI/ML, or a related field is beneficial but not required.
- Continuous, demonstrable learning in AI through publications, open-source contributions, conference speaking, or lab-based experimentation.
- NVIDIA certifications, NVIDIA Deep Learning Institute credentials, or similar are strongly preferred.
- At least one cloud AI/ML certification is preferred.
- Kubernetes certification such as CKA or CKAD is preferred.
- TOGAF 9/10 or equivalent enterprise architecture certification is beneficial.
- Own end-to-end technical delivery of AI/ML engagements, from architecture definition through build oversight and go-live validation.
- Host and chair Architecture Review Board and Technical Design Authority sessions for AI engagements.
- Architect AI infrastructure spanning GPU/accelerator compute, high-performance interconnects, parallel/high-throughput storage, and orchestration.
- Design the AI software stack, including training and fine-tuning pipelines, distributed training frameworks, inference platforms, MLOps/LLMOps tooling, vector databases, and RAG and agentic architectures.
- Define AI governance frameworks covering model risk management, responsible AI, data lineage, bias/fairness testing, explainability, and regulatory alignment.
- Act as a trusted technical advisor to client CTOs, CIOs, and Heads of Data/AI on platform strategy, build-vs-buy decisions, and AI operating model design.
- Lead technical workshops, architecture design sessions, and proof-of-concept builds with cross-functional engineering, data science, and security teams.
- Serve as the technical escalation point for delivery teams and unblock design and implementation issues under time pressure.
- Mentor other architects and engineers on AI systems design.
- Partner with sales and pre-sales to scope AI solutions, size infrastructure, and validate technical feasibility of proposed architectures.
- Define automation, orchestration, and observability standards across the AI stack, from GPU cluster provisioning through to model monitoring in production.
- Architect integration points connecting AI platforms to enterprise networks, third-party systems, and external or service-provider-hosted environments.
- Track the fast-moving AI landscape, including new model architectures, silicon, frameworks, and regulation, and translate relevant developments into our delivery methodology and client recommendations.
- AI
- Airflow
- API
- Ansible
- Architect
- ArgoCD
- CI/CD
- Cloud
- Docker
- Ethernet
- Fabric
- Fine-tuning
- GitOps
- Grafana
- Hardware
- InfiniBand
- Kubeflow
- Kubernetes
- LLM
- Langfuse
- Liquid
- MLflow
- MLOps
- Model Training
- Network
- Prometheus
- PyTorch
- RAG
- REST
- Security
- TOGAF
- TensorFlow
- Terraform
- gRPC
- vLLM
- NodeJS
- AWS
- OpenSearch
- Azure
- ELK
- ETL
- GCP
- Kafka
- Machine Learning
- OpenTelemetry
- OpenShift
- OpenStack
- Python
- VMware
More:
We are World Wide Technology, and we are looking for a deeply technical Enterprise Architect to own the delivery of AI projects end to end, from the silicon and data center design that underpins AI workloads through to the software, MLOps stack, and governance frameworks that make AI trustworthy and defensible at scale. This is a technical hardware-and-software architect role rather than a strategy-only position. The successful candidate will operate across GPU infrastructure, high-performance networking, model training and inference pipelines, and AI risk and governance disciplines. We will have this person lead technical delivery teams for client engagements as the single point of technical accountability from design through go-live, while mentoring delivery teams and shaping our broader AI point of view. The role is Enterprise Architect — Artificial Intelligence within our AI/ML Infrastructure, Platform & Governance practice area, reporting to the Manager, ES&AL, at the Enterprise Architect level. The role is remote with client site travel as required, with travel up to 30% depending on the engagement portfolio, and it is salaried and full-time in Delivery Engineering.
last updated 36 week of 2026
Skills
- Agentic AI
- AI
- Airflow
- Ansible
- API
- Argo CD
- Automation
- AWS
- Azure
- CI/CD
- Cloud
- Data Lineage
- Data Science
- Deep Learning
- DeepSpeed
- Distributed Training
- Docker
- ELK
- Ethernet
- ETL
- Fine Tuning
- GCP
- Generative AI
- GitOps
- Grafana
- gRPC
- HPC
- Infrastructure as Code
- Kafka
- Kubeflow
- Kubernetes
- LLM
- LLMOps
- Machine Learning
- MLflow
- MLOps
- Networking
- Nist
- Node.js
- Observability
- OpenSearch
- OpenShift
- OpenStack
- OpenTelemetry
- Pre-sales
- Prometheus
- Proof of Concept
- Python
- PyTorch
- TensorFlow
- TensorRT
- Terraform
- TLS
- Triton
- Vector Databases
- vLLM
- VMware