HPC Engineer
We are partnering with a industry leader in Singapore to hire experienced HPC Engineers. This role offers the opportunity to support large-scale High Performance Computing (HPC) environments that enable advanced research, AI, data analytics, engineering simulations, and other compute-intensive workloads.
You will be responsible for the design, deployment, administration, and optimisation of HPC infrastructure, including compute clusters, GPU platforms, parallel storage systems, and high-speed networking technologies.
Key Responsibilities
- Deploy, configure, and support HPC cluster infrastructure
- Administer workload management platforms such as Slurm, PBS Pro, or LSF
- Manage GPU computing environments and accelerator platforms
- Support and optimise parallel file systems including Lustre, BeeGFS, or GPFS/Spectrum Scale
- Maintain high-speed interconnect technologies such as InfiniBand and Omni-Path
- Deploy and manage HPC software stacks, including MPI, compilers, scientific libraries, and container technologies
- Monitor cluster utilisation, performance, storage capacity, and overall system health
- Implement automation and configuration management solutions using tools such as Ansible
- Support researchers and technical users in workload optimisation and resource utilisation
- Maintain technical documentation, operational procedures, and capacity planning records
Requirements
HPC Engineer (L2)
- 2-5 years of experience in Linux infrastructure, systems administration, HPC, or related environments
- Experience supporting large-scale Linux platforms and enterprise infrastructure
- Familiarity with job schedulers, storage systems, and networking concepts
- Knowledge of scripting and automation tools
Senior HPC Engineer (L3)
- 5-10 years of experience in HPC or large-scale distributed computing environments
- Strong expertise in cluster administration, performance tuning, and capacity planning
- Experience managing GPU-based computing environments
- Hands-on knowledge of parallel file systems and high-performance networking technologies
- Proven track record supporting mission-critical infrastructure
Preferred Technical Skills
- Linux (Red Hat, Rocky Linux, CentOS)
- Slurm, PBS Pro, LSF
- NVIDIA GPU Platforms
- MPI (OpenMPI, Intel MPI)
- InfiniBand, Omni-Path
- GPFS/Spectrum Scale, Lustre, BeeGFS
- Lmod, Spack, Singularity, Apptainer
- Ansible and infrastructure automation
- Performance monitoring and optimisation tools