freehire launches on Product Hunt on 26 August.

Follow →

Infrastructure Engineer: GPU Fleet (HPC)

You will architect and own the lifecycle of a high-density H200, B200, and B300 GPU fleet. You will manage firmware validation, bare-metal provisioning, decommissioning, liquid-cooled infrastructure, telemetry, automated remediation, NetBox integration, vendor operations, enterprise infrastructure reviews, and 24/7 incident response. You will also lead technical post-mortems, manage escalations, and support capacity planning for major clients.

Responsibilities

  • Own the end-to-end health and lifecycle of H200, B200, and B300 GPU nodes
  • Lead operational oversight of high-density liquid-cooled environments
  • Monitor CDU health, secondary loop telemetry, and GPU thermals
  • Architect telemetry using Prometheus, Grafana, and NVIDIA DCGM
  • Trigger automated node draining, reboots, and health validation
  • Migrate inventory to NetBox DCIM
  • Build API integrations for asset tracking, IPAM, and cabling
  • Serve as the primary technical interface for facility operators and MSPs
  • Set SLA and KPI compliance standards
  • Lead technical post-mortems and manage cluster-level outage escalations
  • Support enterprise deal cycles with capacity planning and infrastructure reviews
  • Participate in a 24/7 on-call rotation
  • Own fleet availability and incident response

Requirements

  • Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility
  • Deep hands-on experience with H200, B200, or B300 systems
  • Expert knowledge of 400G/800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures
  • Strong Linux internals knowledge
  • Proven proficiency building infrastructure automation using Python or Go
  • Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring
  • Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling
  • Familiarity with NVIDIA Mission Control
  • Expertise in Intel TDX or NVIDIA RIM attestation flows
  • Prior experience as an initial infrastructure hire

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available