Point your AI agent at freehire and let it find you a job.

Get the CLI →

ByteBridge

NewBe an early applicant

AIDC Onsite Operation & Maintenance Engineer (Outsource)

Posted
Discussion

AIDC Onsite Operation & Maintenance Engineer (Outsource)

Employment Type: Onsite

Work Location: AIDC Project Site

Work Schedule: 7×24 Shift Rotation (including night shifts, weekends, and statutory holidays)

Specialization: Onsite O&M for Servers, Networks, and HVAC/Electrical Infrastructure

Responsibilities

1. IT Equipment O&M and Incident Management

  • Perform daily inspections, status monitoring, and basic configuration changes for x86 servers, GPU servers, switches, routers, and storage devices.

  • Independently conduct fault diagnosis, replacement of components (HDDs, PSUs, RAM, NICs, etc.), OS reinstallation/debugging, and storage volume configurations.

  • Execute 7×24 On-call alarm response via DCIM and related monitoring tools:

    • Acknowledge alerts within 15 minutes.

    • Initiate response for P1/P2 incidents within 30 minutes.

    • Resolve incidents, implement temporary workarounds, or escalate within 2 hours in principle.

  • Coordinate with project managers, vendors, and technical teams according to escalation protocols to prevent issue proliferation.

  • Assist with server performance checks, hardware health analyses, and basic GPU status troubleshooting.

2. Data Center Infrastructure Collaboration

  • Collaborate with electrical, HVAC, and facility teams to handle cross-system faults (e.g., server overheating caused by AC failure, power anomalies in rack PDUs).

  • Monitor rack-level temperature, humidity, power loads, and environmental conditions to identify local hotspots, overloads, and power risks promptly.

  • Assist in liquid-cooling system inspections and troubleshooting; report and execute emergency measures upon detecting leaks, abnormal temperatures, flow rates, or pressure.

  • Assist in equipment rack mounting/unmounting, cabling, labeling checks, cable management, and rack adjustments.

3. ITIL Processes & Ticket Management

  • Handle Incident, Service Request, Problem, and Change tickets using ITSM tools like ServiceNow or Jira.

  • Strictly follow SLAs for ticket receipt, diagnosis, resolution, escalation, feedback, and closure to ensure end-to-end traceability.

  • Participate in Change Management; conduct risk assessments and operational pre-checks before hardware expansions, replacements, upgrades, or configuration changes.

  • Provide comprehensive event timelines, activity logs, and preliminary Root Cause Analysis (RCA) materials for major incidents.

4. RMA & Spare Parts Management

  • Oversee faulty device diagnosis, RMA applications, spare part requests, vendor coordination, onsite replacement, and repair tracking.

  • Control the average hardware MTTR (Mean Time to Repair) within 4 hours when spare parts and onsite conditions permit.

  • Maintain the onsite spare parts inventory and accurately track check-ins, check-outs, returns, replacements, and write-offs.

  • Periodically audit safety stock and part validity to maximize spare part availability and turnover efficiency.

5. Documentation & Reporting

  • Author and maintain equipment SOPs, inspection checklists, troubleshooting guidebooks, and emergency response procedures.

  • Build a typical incident knowledge base detailing symptoms, diagnostic steps, solutions, and preventive measures.

  • Submit daily/weekly reports, shift handovers, incident reports, and hardware health metrics as required.

  • Ensure all operations, changes, failures, and asset movements are fully logged.

6. Security, Compliance & Collaboration

  • Strictly abide by data center security, information security, ESD protection, access control, and work authorization policies.

  • Refrain from device operations, configuration changes, photography, data copying, or external disclosures without prior approval.

  • Assist project teams with asset inventories, equipment verifications, client inspections, and audits.

  • Execute seamless shift handovers, accurately passing along pending tasks, risks, alerts, and hardware anomalies.

  • Maintain effective cross-functional communication with server, network, storage, facility, procurement, asset management, and vendor teams.

Qualifications

1. Work Experience

  • 2+ years of experience in Data Center IT equipment or AIDC onsite O&M.

  • Ability to independently diagnose server hardware faults and replace components, backed by solid onsite troubleshooting experience.

  • Experience with GPU clusters, supercomputing centers, or Tier III+ data center projects is preferred.

  • Hands-on experience in large-scale server cluster deployment, acceptance testing, racking, or O&M is preferred.

2. Technical Skills

  • Familiarity with x86 servers, GPU servers, blade servers, and standard hardware architectures.

  • Basic knowledge of networking (TCP/IP, VLANs) and storage configurations (RAID, storage volumes).

  • Proficient in using BMC, IPMI, and common server hardware diagnostic utilities.

  • Working knowledge of Linux commands for log queries, status checks, and basic fault localization.

  • Familiarity with monitoring platforms such as Zabbix, Prometheus, or equivalent tools.

  • Mastery of ITIL Incident, Problem, and Change Management workflows, as well as vendor RMA procedures.

  • Basic scripting skills in Shell or Python are preferred.

  • Understanding the operational characteristics of GPUs, liquid-cooling systems, and high-density racks is preferred.

3. Soft Skills & Attributes

  • Flexibility to work 7×24 shifts, night shifts, weekends, and statutory holidays.

  • Strong security mindset, high sense of responsibility, strong execution, and ability to work under pressure.

  • Meticulous and rigorous approach; strict adherence to tickets, SOPs, and authorized scopes.

  • Clear incident reporting and cross-team communication skills, with the ability to leverage frameworks like 5W2H to gather and convey key information efficiently.

Skills

See also

Industrial Engineering jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available