Point your AI agent at freehire and let it find you a job.

Get the CLI →

Cloud Operations Engineer

NewBe an early applicant

Role Overview

We are seeking a seasoned Major Incident & Problem Manager to lead high-priority technology incident resolution, post-incident investigations, and operational service continuity across our core banking and infrastructure platforms. You will direct technical command bridges, establish clear operational accountability during crisis events, and ensure rapid service restoration in compliance with regulatory and ITIL standards.

Key Responsibilities

  • Major Incident Command & Control: Take end-to-end ownership of critical incidents (Sev 1 / Sev 2), coordinating across cross-functional engineering, infrastructure, and application teams to minimize Mean Time to Restore (MTTR).
  • Crisis Communication & Governance: Manage executive escalations, provide concise real-time situation updates to senior leadership and business stakeholders, and ensure full compliance with group technology standards.
  • Problem Management & RCA: Facilitate post-incident reviews using structured analysis methodologies (e.g., 5 Whys, Fishbone) to identify underlying causes, eliminate repeat incidents, and track preventative actions to closure.
  • Operational Reporting & Metrics: Monitor incident patterns, compile KPI dashboards (MTTR, SLA compliance, resolution timelines), and support audit and regulatory reporting deliverables.
  • Continuous Service Improvement: Partner with Command Center and Infrastructure units to enhance automated alerting, runbook execution, and incident logging workflows.

Requirements

  • Bachelor’s degree in Computer Science or equivalent with around 8 years of relevant experience
  • Proven track record leading Major Incident Management (MIM) and Problem Management within enterprise, high-availability IT environments.
  • Strong working knowledge of ITIL service frameworks (ITIL certification required).
  • Working familiarity with enterprise service management platforms (e.g., BMC Remedy, BMC Helix, ServiceNow).
  • Broad technical literacy across enterprise environments: Application Support, End-of-Day (EOD) batch scheduling, infrastructure components (Linux/Unix, Storage, Network), middleware, and transactional workflows (e.g., payment channels).
  • Exceptional crisis communication and stakeholder management skills, with the ability to maintain composure, command technical calls, and align multiple teams under pressure.
  • Strong analytical capabilities in data reporting and presentation using MS Excel (including macros) and PowerPoint.
  • Flexibility to support critical operational escalations or high-impact release windows when required.

See also

DevOps jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available