Senior Data Center Infrastructure & Systems Engineer
Summary
Senior engineer who deploys and runs high-density GPU AI clusters in a Johor Bahru data center: bare-metal OS provisioning, InfiniBand/RoCEv2 network fabric setup, GPU diagnostics and burn-in testing, then ongoing monitoring, troubleshooting, and SLA-driven maintenance. Core stack: Linux (Ubuntu/RHEL), CUDA, BGP/VXLAN, IPMI/Redfish.
We are looking for a Senior Data Center Infrastructure & Systems Engineer to lead the technical deployment, network integration, and diagnostic testing of high-density GPU AI cluster infrastructure. The role covers hardware provisioning, network fabric configuration, and rigorous performance validation, transitioning into ongoing monitoring, troubleshooting, and maintenance of the cluster fleet.
Key Responsibilities
- Configure Out-of-Band (OOB) management networks and execute bare-metal OS provisioning, kernel tuning, and CUDA/driver installations across server nodes.
- Deploy and validate high-speed network fabrics (InfiniBand, RoCEv2 Ethernet), including switch configuration, BGP routing, and VXLAN overlays.
- Perform multi-tier hardware validation, including GPU diagnostics, memory bandwidth benchmarks, power stress tests, and multi-node performance benchmarking.
- Conduct burn-in testing, monitor thermal thresholds, and support functional performance acceptance sign-off.
- Execute SLA-driven work orders including hardware troubleshooting, component swaps, and rack-level maintenance.
- Monitor cluster health telemetry (power, temperature, ECC errors, network performance) to proactively identify hardware issues.
- Respond to infrastructure incidents, isolate faulty hardware, and support non-disruptive maintenance and firmware updates.
- Maintain accurate asset inventory records and comply with physical security and data confidentiality policies.
Qualifications & Experience
- Bachelor's Degree in Computer Engineering, Computer Science, Network Engineering, Systems Administration, or equivalent practical experience.
- 3-5+ years of hands-on experience in high-performance computing (HPC), hyperscale data centers, or AI cluster infrastructure.
- Proven experience with high-density GPU server hardware and liquid-cooled rack systems.
- Strong Linux systems administration skills (Ubuntu/RHEL), PXE provisioning, and scripting (Bash/Python).
- Expertise in high-speed network fabrics: InfiniBand, RoCEv2, BGP, VXLAN, IPAM.
- Experience with GPU diagnostic tools (NVIDIA DCGM, nvidia-smi, Fabric Manager) and IPMI/Redfish APIs.
- Familiarity with ticketing systems and asset tracking workflows.
- Strong diagnostic and troubleshooting skills for complex hardware/network issues.
- Strict adherence to safety and security protocols.
- Ability to adapt between fast-paced deployment work and structured, SLA-driven operations.
Work Location
You will be based in Johor Bahru, within the Iskandar Puteri area.