Infrastructure Engineer: GPU Fleet (HPC)
You will architect and own the lifecycle of a high-density H200, B200, and B300 GPU fleet. You will manage firmware validation, bare-metal provisioning, decommissioning, liquid-cooled infrastructure, telemetry, automated remediation, NetBox integration, vendor operations, enterprise infrastructure reviews, and 24/7 incident response. You will also lead technical post-mortems, manage escalations, and support capacity planning for major clients.
Responsibilities
- Own the end-to-end health and lifecycle of H200, B200, and B300 GPU nodes
- Lead operational oversight of high-density liquid-cooled environments
- Monitor CDU health, secondary loop telemetry, and GPU thermals
- Architect telemetry using Prometheus, Grafana, and NVIDIA DCGM
- Trigger automated node draining, reboots, and health validation
- Migrate inventory to NetBox DCIM
- Build API integrations for asset tracking, IPAM, and cabling
- Serve as the primary technical interface for facility operators and MSPs
- Set SLA and KPI compliance standards
- Lead technical post-mortems and manage cluster-level outage escalations
- Support enterprise deal cycles with capacity planning and infrastructure reviews
- Participate in a 24/7 on-call rotation
- Own fleet availability and incident response
Requirements
- Extensive experience managing large-scale HPC environments or production GPU fleets at a hyperscaler, neocloud, or top-tier research facility
- Deep hands-on experience with H200, B200, or B300 systems
- Expert knowledge of 400G/800G InfiniBand, ConnectX-7 NDR, ConnectX-8 XDR, NVLink, and NVSwitch architectures
- Strong Linux internals knowledge
- Proven proficiency building infrastructure automation using Python or Go
- Deep experience deploying and scaling DCGM-based telemetry and SNMP-based environmental monitoring
- Direct experience with Direct-to-Chip systems, coolant chemistry management, or immersion cooling
- Familiarity with NVIDIA Mission Control
- Expertise in Intel TDX or NVIDIA RIM attestation flows
- Prior experience as an initial infrastructure hire