Server Engineer(AI Cluster)
- Manage the full lifecycle of data center servers, including design, deployment, performance tuning, and validation, as well as the implementation, operation, and maintenance of HPC/AI clusters.
- Lead the deployment of heterogeneous computing resources such as GPUs and XPUs, optimize system performance and stability, and monitor and maintain server operations.
- Prepare technical documentation, provide customer support, and drive the development of automation scripts using Shell, Python, and Ansible.
Key Requirements
- At least 5 years of relevant experience, with hands‑on expertise in building large‑scale HPC clusters with thousands of GPUs, GPU hardware architecture, and parallel computing technologies such as MPI and OpenMP.
- Strong proficiency in virtualization and containerization technologies, with experience in HPL and NCCL benchmarking and high‑performance file systems such as Lustre and GPFS.
- HPC‑related certifications are preferred.
- Professional working proficiency in English.