Solution Architect Cluster Design
Summary
Design and recommend GPU clusters for AI workloads by analyzing customer needs, producing sizing/topology docs, and validating performance benchmarks with Slurm and Kubernetes.
You engage with customers to understand AI workload profiles and translate them into GPU cluster sizing and topology recommendations. You produce cluster design documents, validate performance benchmarks, align Slurm and Kubernetes with topology requirements, and support commercial negotiations with estimates, lead times, and technical risk assessments.
Responsibilities
- Engage with customers to understand workload profiles
- Translate workloads into cluster sizing and topology recommendations
- Produce cluster design documents
- Define and validate performance benchmarks
- Collaborate with the Sesterce OS team on control plane and scheduling requirements
- Support commercial negotiations with BoM estimates, lead times, and technical risk assessments
Requirements
- 5+ years in HPC or AI infrastructure
- Direct involvement in GPU cluster design or technical pre-sales
- Familiarity with NVIDIA Hopper and Blackwell GPU architectures
- Familiarity with NVLink, NVSwitch, and multi-rail InfiniBand
- Ability to interpret AI workload performance profiles
- Strong written communication
- Background with cloud providers, HPC centers, or AI infrastructure vendors