Point your AI agent at freehire and let it find you a job.

Get the CLI →

Lintasarta

New

Data Center Engineer (L2)

Posted 2 views
Discussion

Summary

Onsite 24×7 L2 data center engineer at Lintasarta in Indonesia, acting as the technical resolution authority for escalated infrastructure incidents in a large-scale GPU cloud — diagnosing and fixing GPU servers, liquid-cooling systems, InfiniBand/converged networking, and storage, and writing RCAs.

The L2 Engineer is the technical resolution authority for all infrastructure faults that exceed L1's first-response capability. Operating onsite 24×7, L2 engineers own each incident ticket from the point of escalation through to hardware-level resolution or formal handoff to L3, maintaining ticket ownership throughout the entire escalation lifecycle. The role demands deep, hands-on expertise across GPU compute, liquid-cooling systems, high-speed network fabric, and server infrastructure in a large-scale, mission-critical GPU cloud environment.

Key Responsibilities:

  • Performs in-depth fault diagnosis and hardware-level resolution across the full infrastructure stack, including GPU server faults (card detection, PCIe/NVLink link status, thermal and power anomalies, ECC error tracking), liquid-cooling system anomalies (CDU and secondary loop: inlet/outlet temperature, flow rate, differential pressure), NVIDIA InfiniBand and converged network faults, and enterprise server and storage issues.
  • Authors Root Cause Analysis (RCA) documentation for every incident, capturing fault timeline, diagnostic findings, actions taken, and preventive recommendations.
  • Governs all infrastructure changes through the CAB, accountable for change request submission, approval coordination, maintenance window execution, rollback planning, and post-change verification.
  • Initiates formal escalation to L3/OEM with a complete evidence package (diagnostic logs, DCGM output, hardware status reports, incident timelines) when standard diagnostic and replacement procedures are exhausted, while retaining ticket ownership until full resolution.
  • Maintains firmware and driver baseline compliance across all in-scope infrastructure, monitoring CVE advisories and OEM bulletins, executing baseline drift corrections within approved maintenance windows, with minimum 90-day configuration backup retention.
  • Supports the knowledge transfer programme through a shadow/pairing model with IOH's embedded operational counterparts from service commencement.

Qualifications & Skills:

  • Min. 3–5 years in Data Center / infrastructure operations, with direct hands-on experience on GPU servers / HPC clusters (not general servers only).
  • In-depth GPU diagnostics (nvidia-smi/lspci, PCIe/NVLink, thermal & power, ECC error tracking).
  • Rack-scale GB200 NVL72 (Schedule A - Voltage).
  • Liquid-cooling systems (CDU & secondary loop), operation, monitoring, and anomaly handling (Schedule A - Voltage).
  • NVIDIA InfiniBand + converged/management networking.
  • Enterprise server & storage: fault isolation & vendor escalation.
  • Strong RCA documentation and technical writing skills.

Skills

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available