HPC Engineer

We are partnering with a industry leader in Singapore to hire experienced HPC Engineers. This role offers the opportunity to support large-scale High Performance Computing (HPC) environments that enable advanced research, AI, data analytics, engineering simulations, and other compute-intensive workloads.

You will be responsible for the design, deployment, administration, and optimisation of HPC infrastructure, including compute clusters, GPU platforms, parallel storage systems, and high-speed networking technologies.

Key Responsibilities

  • Deploy, configure, and support HPC cluster infrastructure
  • Administer workload management platforms such as Slurm, PBS Pro, or LSF
  • Manage GPU computing environments and accelerator platforms
  • Support and optimise parallel file systems including Lustre, BeeGFS, or GPFS/Spectrum Scale
  • Maintain high-speed interconnect technologies such as InfiniBand and Omni-Path
  • Deploy and manage HPC software stacks, including MPI, compilers, scientific libraries, and container technologies
  • Monitor cluster utilisation, performance, storage capacity, and overall system health
  • Implement automation and configuration management solutions using tools such as Ansible
  • Support researchers and technical users in workload optimisation and resource utilisation
  • Maintain technical documentation, operational procedures, and capacity planning records

Requirements

HPC Engineer (L2)

  • 2-5 years of experience in Linux infrastructure, systems administration, HPC, or related environments
  • Experience supporting large-scale Linux platforms and enterprise infrastructure
  • Familiarity with job schedulers, storage systems, and networking concepts
  • Knowledge of scripting and automation tools

Senior HPC Engineer (L3)

  • 5-10 years of experience in HPC or large-scale distributed computing environments
  • Strong expertise in cluster administration, performance tuning, and capacity planning
  • Experience managing GPU-based computing environments
  • Hands-on knowledge of parallel file systems and high-performance networking technologies
  • Proven track record supporting mission-critical infrastructure

Preferred Technical Skills

  • Linux (Red Hat, Rocky Linux, CentOS)
  • Slurm, PBS Pro, LSF
  • NVIDIA GPU Platforms
  • MPI (OpenMPI, Intel MPI)
  • InfiniBand, Omni-Path
  • GPFS/Spectrum Scale, Lustre, BeeGFS
  • Lmod, Spack, Singularity, Apptainer
  • Ansible and infrastructure automation
  • Performance monitoring and optimisation tools