Solutions Architect - AI Factory

Summary

Designs and deploys AI training/inference clusters using NVIDIA GPUs, Kubernetes, and networking tech, advising customers on NVIDIA Reference Architectures.

About

NVIDIA is seeking an outstanding Solutions Architect for AI Factory to assist and support customers that are building solutions with our newest AI technology. At NVIDIA, our Solutions Architects work across different teams and enjoy helping customers with the latest Accelerated Computing and Deep Learning software and hardware platforms. You will become a trusted technical advisor with our customers and work on exciting projects and proof-of-concepts passionate about AI Factory.

Responsibilities

• Maintain an up-to-date understanding of the philosophy, architecture, and deployment methods of various evolving NVIDIA Reference Architectures—e.g., NVIDIA DGX SuperPOD Reference Architecture, NVIDIA Cloud Partner Reference Design, and NVIDIA Enterprise Reference Architecture.
• Analyze and understand the requirements of customer-initiated AI training or inference clusters.
• Identify the NVIDIA Reference Architecture that best matches customer needs and effectively communicate its value proposition to collaborators.
• Facilitate seamless communication between NVIDIA's internal deployment teams and customers during the implementation of AI clusters based on Reference Architectures.
• Provide hands-on technical support to developers after the AI Factory has been deployed, ensuring that AI training and inference workloads run effectively on the infrastructure.

Requirements

• Bachelor's degree or higher in Computer Science, Computer Engineering, or a related technical field.
• Minimum of 7 years of hands-on experience designing, deploying, and operating AI training or inference clusters with server, storage, and network, that leverage Kubernetes with NVIDIA GPUs.
• Foundational knowledge and experience with networking technologies—such as InfiniBand and Ethernet—in AI cluster environments, including compute fabric interconnects between GPU servers, storage fabric integration, and out-of-band and in-band management fabrics.

Preferred

• Proficiency in key technologies including: Container Runtime Interface (CRI), Container Network Interface (CNI), Calico, NVIDIA GPU Operator, NVIDIA Network Operator, and Kubeflow Training Operator.
• Knowledge and experience of facilities being used in building AI Factories, including cooling devices, power distribution, or plumbing.

Benefits

자세한 사항은 홈페이지 참조