IT Operations Engineer (Data Centre & Batch Monitoring)
Summary
On-site IT Operations Engineer in Hong Kong monitoring a business-critical bank data centre on 24/7 rotating shifts: watching batch jobs, dashboards and alerts, doing first-level troubleshooting, logging tickets and escalating incidents. Ops/monitoring experience needed; Windows/Linux, ServiceNow and Control-M-type schedulers are pluses.
Working Arrangement: On-site, 24/7 rotating shifts
Job SummaryWe are seeking an IT Operations Engineer to support the daily operation and monitoring of a business-critical banking data-centre environment.
The successful candidate will monitor systems, scheduled batch activities and operational alerts, perform first-level troubleshooting, maintain accurate records and escalation incidents to the relevant technical teams.
Candidates with experience in data-centre operations, command-centre monitoring, production support, network operations or system operations are encouraged to apply. Training and operational procedures will be provided for candidates who possess relevant 24/7 operations experience but have limited exposure to batch-scheduling tools.
Key Responsibilities- Monitor data-centre systems, operational dashboards, alerts and scheduled batch activities.
- Verify that scheduled jobs and operational tasks are completed successfully and on time.
- Identify job failures, system alerts, delays or other abnormalities.
- Perform first-level troubleshooting and recovery actions according to established procedures.
- Restart or rerun failed activities when authorised and upscale unresolved incidents to the relevant support teams.
- Monitor servers, network devices, applications and infrastructure availability.
- Create, update and follow incident tickets using the designated service-management system.
- Maintain accurate shift logs, incident records, operational checklists and handover reports.
- Coordinate with application, infrastructure, network and system-support teams during incidents.
- Support scheduled maintenance, system changes, backup activities and data-centre operational tasks.
- Prepare regular operational and incident-status reports.
- Follow the bank’s security, compliance, change-management and operational-control procedures.
- Identify recurring operational issues and recommend improvements where appropriate.
- Experience in one or more of the following areas:
- IT command-centre or control-room monitoring
- Production or application support
- System or infrastructure operations
- Batch-job monitoring or scheduling
- Experience supporting a 24/7 production or business-critical environment is preferred.
- Familiarity with Windows, Unix/Linux, servers, networks or enterprise infrastructure.
- Basic troubleshooting and incident-escalation experience.
- Experience using ticketing tools such as ServiceNow would be advantageous.
- Exposure to Control-M, AutoSys, IBM/HCL Workload Scheduler or similar tools is strongly preferred but not mandatory.
- Experience in banking, financial services or another regulated environment would be advantageous.
- Able to follow detailed operating procedures and maintain accurate records.
- Responsible, attentive and able to remain composed when responding to operational incidents.
- Good communication and teamwork skills.
- Good command of Cantonese and English; Mandarin would be advantageous.