Point your AI agent at freehire and let it find you a job.

Get the CLI →

Data Center L2 Hardware & GPU Debug Engineer

Discussion

Summary

Debug and repair enterprise x86 servers and AI/GPU platforms, analyze logs via BMC/IPMI/Redfish, and write RCA reports in a mission-critical data center environment.

Are you passionate about enterprise server hardware, x86 architecture, and cutting-edge AI / GPU platforms? We are looking for an experienced L2 Hardware & Failure Analysis Engineer to join our team supporting mission-critical environments!

Language: Fluent English (Mandatory for global escalations)

  • Perform hands-on Break/Fix and L10/L11 failure analysis on enterprise x86 servers and AI/GPU platforms (CPU, DDR5, Motherboards, PSUs, PCIe).
  • Extract and analyze low-level event logs using BMC, IPMI, and Redfish interfaces.
  • Diagnose network & storage controller issues (PCIe, iSCSI, RoCE, SAS, Fibre Channel).
  • Execute post-repair validations, firmware flashing (BIOS/BMC), and quality checks in production/lab environments.
  • Write Root Cause Analysis (RCA) reports, SOPs, and KB articles.
  • 5+ years of hands-on experience in enterprise server hardware troubleshooting, validation, or Failure Analysis (FA).
  • Deep knowledge of x86 architecture, Linux OS command line, and server management.
  • Direct experience with GPU Servers (NVIDIA / AMD platforms) is a massive plus!
  • Experience in companies like Intel, Foxconn, Ingrasys, Jabil, Wistron, Oracle, Microsoft, or hyperscale Data Centers.

Skills

See also

Hardware jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available