freehire launches on Product Hunt on 26 August.

Follow →

Data Center Site Manager

You manage daily data center operations and lead six shift Operations Engineers. You ensure continuous infrastructure availability, oversee AI and HPC systems, coordinate incident resolution and maintenance, manage staffing and vendors, monitor operational performance, and provide hands-on support during critical events.

Responsibilities

  • Lead daily data center site operations
  • Supervise and manage six shift Operations Engineers
  • Plan manpower and schedule shifts
  • Assign tasks and manage performance
  • Ensure 24x7 operational coverage
  • Escalate incidents and coordinate issue resolution
  • Oversee AI and HPC infrastructure operation and maintenance
  • Establish and improve SOPs, EOPs, and preventive maintenance programs
  • Monitor site health, KPIs, incidents, and infrastructure performance
  • Coordinate hardware installation, rack and stack activities, commissioning, expansion, and lifecycle management
  • Review and approve maintenance activities, change requests, incident reports, and handover records
  • Ensure policy, safety, security, and operational compliance
  • Coordinate with engineering, network, facilities, and vendor teams
  • Participate in on-call duties and provide hands-on operational support

Requirements

  • Bachelor's degree or above
  • Data center operations experience
  • IT infrastructure experience
  • HPC or AI infrastructure management experience
  • Minimum 5 years of relevant experience
  • Minimum 2 years of team leadership or people management experience
  • 24x7 shift operations
  • NVIDIA GB200 and GB300 cluster knowledge
  • GPU server knowledge
  • x86 server knowledge
  • Storage system knowledge
  • Ethernet networking
  • InfiniBand networking
  • NVIDIA GPU architecture
  • NVLink
  • NVSwitch
  • Server hardware troubleshooting
  • Firmware management
  • Hardware lifecycle management
  • Structured cabling
  • Linux administration
  • System and service management
  • Hardware and performance diagnostics
  • Log analysis
  • Network troubleshooting
  • Scripting and automation
  • Shift scheduling
  • Incident and escalation management
  • Performance management
  • SOP and EOP development
  • Vendor coordination
  • Communication
  • Decision-making

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available