GPU AI/HPC DevOps Engineer: Build & Auto-Scale Clusters
Summary
The DevOps Engineer will design, deploy, and manage GPUaaS infrastructure for AI and HPC workloads across on-prem and cloud environments. The role focuses on automation, monitoring, and building CI/CD pipelines using technologies like Kubernetes, Slurm, and GPU drivers.
Singtel is seeking a DevOps Engineer to design, deploy and operate GPUaaS infrastructure for AI/HPC workloads. The role spans on-prem and cloud resources, with a focus on automation, monitoring, and robust CI/CD pipelines for GPU-accelerated applications.
You will work with Slurm, Kubernetes, and GPU drivers, contributing to performance optimizations and scalable, secure multi-tenant environments. Strong Linux skills and scripting are essential.