AI/GPU Platform Resident Engineer (Data Center)
Posted
- Provide operational support for the GB300 NVL72 AI platform.
- Support PowerEdge management node operations and administration.
- Perform NVIDIA platform testing, validation, and performance troubleshooting.
Requirements
- Strong understanding of GPU-based AI infrastructure and rack-scale computing platforms.
- Hands-on experience with NVIDIA platform validation, health checks, and post-installation testing.
- Knowledge of firmware management, driver lifecycle, operating system baseline control, and upgrade sequencing.
- Ability to perform performance analysis and troubleshoot issues across GPU, CPU, memory, and I/O layers.
- Familiarity with cluster bring-up, node acceptance testing, and operational readiness validation.
- Working knowledge of Linux administration and hardware-level diagnostics.
- Ability to identify and correlate platform issues across storage, networking, and compute infrastructure.