Data Center Engineer (L2)
Summary
Onsite 24×7 L2 data center engineer at Lintasarta in Indonesia, acting as the technical resolution authority for escalated infrastructure incidents in a large-scale GPU cloud — diagnosing and fixing GPU servers, liquid-cooling systems, InfiniBand/converged networking, and storage, and writing RCAs.
The L2 Engineer is the technical resolution authority for all infrastructure faults that exceed L1's first-response capability. Operating onsite 24×7, L2 engineers own each incident ticket from the point of escalation through to hardware-level resolution or formal handoff to L3, maintaining ticket ownership throughout the entire escalation lifecycle. The role demands deep, hands-on expertise across GPU compute, liquid-cooling systems, high-speed network fabric, and server infrastructure in a large-scale, mission-critical GPU cloud environment.
Key Responsibilities:
- Performs in-depth fault diagnosis and hardware-level resolution across the full infrastructure stack, including GPU server faults (card detection, PCIe/NVLink link status, thermal and power anomalies, ECC error tracking), liquid-cooling system anomalies (CDU and secondary loop: inlet/outlet temperature, flow rate, differential pressure), NVIDIA InfiniBand and converged network faults, and enterprise server and storage issues.
- Authors Root Cause Analysis (RCA) documentation for every incident, capturing fault timeline, diagnostic findings, actions taken, and preventive recommendations.
- Governs all infrastructure changes through the CAB, accountable for change request submission, approval coordination, maintenance window execution, rollback planning, and post-change verification.
- Initiates formal escalation to L3/OEM with a complete evidence package (diagnostic logs, DCGM output, hardware status reports, incident timelines) when standard diagnostic and replacement procedures are exhausted, while retaining ticket ownership until full resolution.
- Maintains firmware and driver baseline compliance across all in-scope infrastructure, monitoring CVE advisories and OEM bulletins, executing baseline drift corrections within approved maintenance windows, with minimum 90-day configuration backup retention.
- Supports the knowledge transfer programme through a shadow/pairing model with IOH's embedded operational counterparts from service commencement.
Qualifications & Skills:
- Min. 3–5 years in Data Center / infrastructure operations, with direct hands-on experience on GPU servers / HPC clusters (not general servers only).
- In-depth GPU diagnostics (nvidia-smi/lspci, PCIe/NVLink, thermal & power, ECC error tracking).
- Rack-scale GB200 NVL72 (Schedule A - Voltage).
- Liquid-cooling systems (CDU & secondary loop), operation, monitoring, and anomaly handling (Schedule A - Voltage).
- NVIDIA InfiniBand + converged/management networking.
- Enterprise server & storage: fault isolation & vendor escalation.
- Strong RCA documentation and technical writing skills.