Point your AI agent at freehire and let it find you a job.

Get the CLI →

Genesis Networks Pte Ltd

Senior Data Center Infrastructure & Systems Engineer

Posted 2 views
Discussion

Summary

Senior engineer who deploys and runs high-density GPU AI clusters in a Johor Bahru data center: bare-metal OS provisioning, InfiniBand/RoCEv2 network fabric setup, GPU diagnostics and burn-in testing, then ongoing monitoring, troubleshooting, and SLA-driven maintenance. Core stack: Linux (Ubuntu/RHEL), CUDA, BGP/VXLAN, IPMI/Redfish.

We are looking for a Senior Data Center Infrastructure & Systems Engineer to lead the technical deployment, network integration, and diagnostic testing of high-density GPU AI cluster infrastructure. The role covers hardware provisioning, network fabric configuration, and rigorous performance validation, transitioning into ongoing monitoring, troubleshooting, and maintenance of the cluster fleet.

Key Responsibilities

  • Configure Out-of-Band (OOB) management networks and execute bare-metal OS provisioning, kernel tuning, and CUDA/driver installations across server nodes.
  • Deploy and validate high-speed network fabrics (InfiniBand, RoCEv2 Ethernet), including switch configuration, BGP routing, and VXLAN overlays.
  • Perform multi-tier hardware validation, including GPU diagnostics, memory bandwidth benchmarks, power stress tests, and multi-node performance benchmarking.
  • Conduct burn-in testing, monitor thermal thresholds, and support functional performance acceptance sign-off.
  • Execute SLA-driven work orders including hardware troubleshooting, component swaps, and rack-level maintenance.
  • Monitor cluster health telemetry (power, temperature, ECC errors, network performance) to proactively identify hardware issues.
  • Respond to infrastructure incidents, isolate faulty hardware, and support non-disruptive maintenance and firmware updates.
  • Maintain accurate asset inventory records and comply with physical security and data confidentiality policies.

Qualifications & Experience

  • Bachelor's Degree in Computer Engineering, Computer Science, Network Engineering, Systems Administration, or equivalent practical experience.
  • 3-5+ years of hands-on experience in high-performance computing (HPC), hyperscale data centers, or AI cluster infrastructure.
  • Proven experience with high-density GPU server hardware and liquid-cooled rack systems.
  • Strong Linux systems administration skills (Ubuntu/RHEL), PXE provisioning, and scripting (Bash/Python).
  • Expertise in high-speed network fabrics: InfiniBand, RoCEv2, BGP, VXLAN, IPAM.
  • Experience with GPU diagnostic tools (NVIDIA DCGM, nvidia-smi, Fabric Manager) and IPMI/Redfish APIs.
  • Familiarity with ticketing systems and asset tracking workflows.
  • Strong diagnostic and troubleshooting skills for complex hardware/network issues.
  • Strict adherence to safety and security protocols.
  • Ability to adapt between fast-paced deployment work and structured, SLA-driven operations.

Work Location

You will be based in Johor Bahru, within the Iskandar Puteri area.

Skills

What Senior Software Engineering jobs ask for — and how much of it you have →
Apply

See also

Software Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available