L3 Data Center Engineer
Summary
Top-tier (L3) data center engineer acting as the final escalation authority for GPU/HPC infrastructure incidents: firmware-level diagnostics on NVIDIA GPU servers, InfiniBand fabric and liquid-cooling troubleshooting, OEM escalation ownership, and mentoring of L1/L2 staff, on a 24x7 on-call rotation.
The L3 Engineer / GPU Infrastructure Specialist is the highest technical authority within the managed service delivery team, engaged when L2 has exhausted all standard diagnostic and replacement procedures and the incident requires firmware-level analysis, OEM-specific tooling, or direct vendor engineering support. Operating on a 24×7 on-call basis with the ability to mobilise onsite when required, L3 specialists serve simultaneously as deep technical resolvers, OEM relationship owners, and knowledge anchors for the broader delivery team.
Responsibilities:
- Handles incidents beyond L2 resolution scope, including firmware-level failure analysis, NVLink/PCIe fabric anomalies, deep XID and ECC error investigations, and the application of OEM-specific diagnostic tooling unavailable at lower tiers; produces formal RCA and failure reports for every L3 engagement.
- Owns the complete OEM escalation lifecycle, from defining trigger conditions and assembling evidence packages, to direct interface with NVIDIA, server OEM, and network OEM engineering teams, through to resolution tracking and formal closure.
- Escalates to the Principal Expert when incidents exceed L3 resolution capability, including cases requiring manufacturer-level engineering intervention, multi-vendor cross-stack failures, or issues with no established resolution precedent.
- Serves as the authoritative technical reference point for L1 and L2 personnel throughout incident handling, providing real-time diagnostic guidance and directing cross-stack resolution strategy.
- Leads advanced troubleshooting of the NVIDIA InfiniBand backend fabric (fabric health, NCCL collective performance, network congestion) and root-cause analysis of liquid-cooling system failures at CDU and secondary loop level.
- Attends regular governance review meetings to advise on failure trends, troubleshooting strategy, and infrastructure risk.
- Contributes to the operational knowledge base through escalation playbooks, known-issue documentation, and structured competency uplift of the L1/L2 team via the shadow/pairing model.
Qualifications & Skills:
- Min. 7 years in Data Center / infrastructure, with deep hands-on expertise on GPU servers / HPC clusters at firmware and OEM level, beyond day-to-day operations.
- Advanced GPU diagnostics (Nvidia-smi/lspci, PCIe/NVLink link & topology, thermal/power, ECC & XID error tracking).
- NVIDIA InfiniBand networking, fabric troubleshooting, link/subnet manager, OEM escalation.
- Liquid cooling systems (CDU & secondary loop) to root-cause analysis level.
- Demonstrated end-to-end OEM escalation leadership.
- Manufacturer-level capability for complex, cross-stack infrastructure issues.