SRE L1 Support/Cloud Platform Ops Engineers

You monitor GPU clusters, networks, storage systems, and environmental sensors; respond to alerts and execute incident runbooks; triage and replace hardware; perform standard remediation; collect diagnostics for escalation; manage incident tickets; complete physical data center tasks; conduct shift handoffs; maintain runbooks; and assist with hardware deployment, firmware updates, and inventory management.

Responsibilities

  • Monitor GPU cluster health, network status, storage systems, and environmental sensors
  • Respond to alerts and execute runbooks for GPU, network, node, and storage incidents
  • Identify failed GPUs, NICs, PSUs, disks, and cables
  • Execute GPU resets, node drains and reboots, link reseating, and BMC recovery
  • Collect logs, DCGM output, network diagnostics, and hardware health reports for escalation
  • Manage incident tickets through resolution or escalation
  • Perform cable installation, hardware swap-outs, rack and stack, and labeling
  • Execute shift handoffs with the APAC operations team
  • Maintain and update operational runbooks
  • Assist with hardware deployment, firmware updates, and inventory management

Requirements

  • 2+ years of experience in NOC, data center operations, or IT support
  • Basic Linux system administration
  • Familiarity with Prometheus, Grafana, Nagios, or equivalent monitoring tools
  • Experience with ServiceNow or Jira Service Management
  • Ability to perform rack and stack, cabling, and hardware replacement
  • Strong communication skills
  • Ability to work 8AM-8PM PST shifts with rotation

Benefits

  • Attractive welfare benefits

See also

SRE jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available