Point your AI agent at freehire and let it find you a job.

Get the CLI →

Lintasarta

New

Data Center Engineer (L1)

Posted Updated 2 views
Discussion

Summary

L1 Data Center Engineer at Lintasarta providing 24×7 onsite coverage across NOC surveillance monitoring and hands-on data center floor response for GPU infrastructure. Day to day involves monitoring dashboards (DCGM, NetQ, UFM, Grafana, ServiceNow), classifying and escalating incidents, performing physical rack inspections, and supporting preventive maintenance.

The L1 Engineer forms the foundational layer of the tiered support model, operating across two complementary sub-functions within the same 24×7 onsite coverage model: Surveillance (NOC-based continuous monitoring) and Data Center (onsite physical response). Together, these two sub-functions ensure that every infrastructure event, whether detected remotely through dashboards or observed directly on the data center floor, is captured, classified, and escalated with precision and speed. L1 is the first point of contact for all incidents and the initiator of the escalation path to L2 and L3.


Responsibilities:



  • Maintains continuous 24×7 monitoring of all GPU infrastructure dashboards, covering compute health, GPU utilization, fabric and network status, power and cooling parameters, and environmental conditions, using platforms including DCGM, NetQ, UFM, Grafana, and ServiceNow.

  • Classifies all alarms by severity (P1–P4), validates against false-positive filters, and dispatches through the appropriate escalation workflow with full SLA tracking.

  • Creates accurate, complete incident tickets in the ITSM platform and dispatches to the appropriate tier (L1 DC for physical check, L2 for technical diagnosis, or L3 for complex escalation).

  • Provides P1 status updates every 30 minutes until resolution; monitors SLA countdown for all active tickets and proactively escalates tickets at risk of breach.

  • Conducts daily synthetic health checks: canary jobs, NCCL bandwidth tests, fabric monitoring, and log pipeline health verification.

  • Conducts regular physical walkthroughs and rack inspections, verifying LED indicators, cabling integrity, power supply status, and the physical condition of GPU servers, NVLink switches, and CDUs.

  • Responds to ticket dispatch from L1 Surveillance for direct on-floor physical checks, and performs basic hardware verification and initial corrective actions (e.g., reboot/power-cycle via BMC) before escalating to L2.

  • Executes Emergency Response Procedures (EPO or Loop Isolation protocols) within 15 minutes of alert confirmation for P1 environmental incidents including cooling failures, power irregularities, or liquid coolant leaks.

  • Monitors liquid-cooling system parameters (CDU and secondary loop) and coordinates with the DC facilities team on environmental anomalies.

  • Supports preventive maintenance activities: rack deep cleaning (quarterly), node health sweep (monthly), and cold-spare rotation.

  • Executes a structured shift handover at every transition, ensuring full situational awareness of active incidents, open tickets, pending dispatches, and infrastructure anomalies is formally transferred to the incoming shift with zero information loss.


Qualifications & Skills :



  • Min. 1–2 years in Data Center / NOC / IT Operations.

  • Basic TCP/IP networking.

  • Proficient with monitoring tools (Grafana / Zabbix / DCIM / NMS) & ITSM ticketing systems.

  • Solid understanding of SLA concepts, incident prioritization, and escalation workflows.

Skills

Apply

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available