AIDC Onsite Operation & Maintenance Engineer (Outsource)
AIDC Onsite Operation & Maintenance Engineer (Outsource)
Employment Type: Onsite
Work Location: AIDC Project Site
Work Schedule: 7×24 Shift Rotation (including night shifts, weekends, and statutory holidays)
Specialization: Onsite O&M for Servers, Networks, and HVAC/Electrical Infrastructure
Responsibilities
1. IT Equipment O&M and Incident Management
Perform daily inspections, status monitoring, and basic configuration changes for x86 servers, GPU servers, switches, routers, and storage devices.
Independently conduct fault diagnosis, replacement of components (HDDs, PSUs, RAM, NICs, etc.), OS reinstallation/debugging, and storage volume configurations.
Execute 7×24 On-call alarm response via DCIM and related monitoring tools:
Acknowledge alerts within 15 minutes.
Initiate response for P1/P2 incidents within 30 minutes.
Resolve incidents, implement temporary workarounds, or escalate within 2 hours in principle.
Coordinate with project managers, vendors, and technical teams according to escalation protocols to prevent issue proliferation.
Assist with server performance checks, hardware health analyses, and basic GPU status troubleshooting.
2. Data Center Infrastructure Collaboration
Collaborate with electrical, HVAC, and facility teams to handle cross-system faults (e.g., server overheating caused by AC failure, power anomalies in rack PDUs).
Monitor rack-level temperature, humidity, power loads, and environmental conditions to identify local hotspots, overloads, and power risks promptly.
Assist in liquid-cooling system inspections and troubleshooting; report and execute emergency measures upon detecting leaks, abnormal temperatures, flow rates, or pressure.
Assist in equipment rack mounting/unmounting, cabling, labeling checks, cable management, and rack adjustments.
3. ITIL Processes & Ticket Management
Handle Incident, Service Request, Problem, and Change tickets using ITSM tools like ServiceNow or Jira.
Strictly follow SLAs for ticket receipt, diagnosis, resolution, escalation, feedback, and closure to ensure end-to-end traceability.
Participate in Change Management; conduct risk assessments and operational pre-checks before hardware expansions, replacements, upgrades, or configuration changes.
Provide comprehensive event timelines, activity logs, and preliminary Root Cause Analysis (RCA) materials for major incidents.
4. RMA & Spare Parts Management
Oversee faulty device diagnosis, RMA applications, spare part requests, vendor coordination, onsite replacement, and repair tracking.
Control the average hardware MTTR (Mean Time to Repair) within 4 hours when spare parts and onsite conditions permit.
Maintain the onsite spare parts inventory and accurately track check-ins, check-outs, returns, replacements, and write-offs.
Periodically audit safety stock and part validity to maximize spare part availability and turnover efficiency.
5. Documentation & Reporting
Author and maintain equipment SOPs, inspection checklists, troubleshooting guidebooks, and emergency response procedures.
Build a typical incident knowledge base detailing symptoms, diagnostic steps, solutions, and preventive measures.
Submit daily/weekly reports, shift handovers, incident reports, and hardware health metrics as required.
Ensure all operations, changes, failures, and asset movements are fully logged.
6. Security, Compliance & Collaboration
Strictly abide by data center security, information security, ESD protection, access control, and work authorization policies.
Refrain from device operations, configuration changes, photography, data copying, or external disclosures without prior approval.
Assist project teams with asset inventories, equipment verifications, client inspections, and audits.
Execute seamless shift handovers, accurately passing along pending tasks, risks, alerts, and hardware anomalies.
Maintain effective cross-functional communication with server, network, storage, facility, procurement, asset management, and vendor teams.
Qualifications
1. Work Experience
2+ years of experience in Data Center IT equipment or AIDC onsite O&M.
Ability to independently diagnose server hardware faults and replace components, backed by solid onsite troubleshooting experience.
Experience with GPU clusters, supercomputing centers, or Tier III+ data center projects is preferred.
Hands-on experience in large-scale server cluster deployment, acceptance testing, racking, or O&M is preferred.
2. Technical Skills
Familiarity with x86 servers, GPU servers, blade servers, and standard hardware architectures.
Basic knowledge of networking (TCP/IP, VLANs) and storage configurations (RAID, storage volumes).
Proficient in using BMC, IPMI, and common server hardware diagnostic utilities.
Working knowledge of Linux commands for log queries, status checks, and basic fault localization.
Familiarity with monitoring platforms such as Zabbix, Prometheus, or equivalent tools.
Mastery of ITIL Incident, Problem, and Change Management workflows, as well as vendor RMA procedures.
Basic scripting skills in Shell or Python are preferred.
Understanding the operational characteristics of GPUs, liquid-cooling systems, and high-density racks is preferred.
3. Soft Skills & Attributes
Flexibility to work 7×24 shifts, night shifts, weekends, and statutory holidays.
Strong security mindset, high sense of responsibility, strong execution, and ability to work under pressure.
Meticulous and rigorous approach; strict adherence to tickets, SOPs, and authorized scopes.
Clear incident reporting and cross-team communication skills, with the ability to leverage frameworks like 5W2H to gather and convey key information efficiently.