NOC & Application Monitoring Manager
Envision Employment Solutions is currently looking for a NOC & Application Monitoring Manager for one of our partners, a leading Digital Bank!
THE ROLE
You lead the NOC and Application Monitoring function, providing 24x7 operational oversight across the digital bank and serves as the first line of response for technology incidents. The team monitors applications, infrastructure and critical customer journeys, responds to alerts, performs initial triage and remediation, and escalates incidents to the appropriate technical teams.
The manager is responsible for the effectiveness of the function, including staffing and shift coverage, monitoring standards, alert quality, escalation paths, incident procedures, operational reporting and readiness for major incidents. The role also drives proactive identification of performance, capacity and availability risks before they become customer-impacting incidents.
WHAT YOU’LL DO
- Lead the 24x7 NOC and Application Monitoring team across production systems, applications, networks and infrastructure.
- Provide L1 monitoring, triage and first-line response using service-specific dashboards, alerts, playbooks and escalation paths.
- Stand up observability and Site Reliability Engineering practice, and drive root cause analysis and continuous improvement.
- Coordinate operational incidents across engineering, infrastructure, application, security and 3rd parties.
- Define monitoring, alerting, escalation & operational-readiness standards for service teams to follow.
- Ensure service owners maintain effective dashboards, actionable alerts, current playbooks and clear support ownership.
- Review monitoring gaps, alert quality, recurring incidents and service-level breaches, and drive corrective actions with accountable teams.
- Maintain the enterprise view of availability, performance, capacity and critical customer journeys.
- Coordinate major incident readiness, shift handovers, operational reporting and continuous improvement of NOC processes.
WHAT WE’RE LOOKING FOR
- 7+ years in NOC operations, application monitoring, service operations or a related field.
- Experience leading 24x7 teams and coordinating incidents across multiple technical functions.
- Strong technical knowledge of observability platforms such as Elastic, Splunk, Dynatrace, Prometheus or Grafana.
- A working command of ITIL incident, event and problem management.
- Experience defining operational standards, escalation models, support processes and service-level reporting.
- The ability to challenge monitoring quality without taking ownership away from engineering.
- Financial-services experience is valuable but not essential.
YOU’LL THRIVE HERE IF YOU
- You believe an incident a customer reports first is a monitoring failure, not bad luck
- You stay calm at three in the morning and let the runbook do its work
- You read a health report and see the problem forming before it breaks
As published by workable
First name, Last name, Email, Headline, Phone, Address, Photo, Education, Experience, Summary, Resume, Cover letter
- How did you hear about us? choose one
- If you were referred to this role by a recruiter, please provide their name. written answer
- Do you have a minimum of 7 years of experience in NOC operations, application monitoring, or Site Reliability Engineering? yes / no
- Have you worked in banking, financial services, or another regulated industry? yes / no
- Do you have experience managing a 24/7 shift based monitoring or NOC team? yes / no
- Do you have hands on experience with observability or monitoring tools (e.g., Splunk, Datadog, Dynatrace, Prometheus/Grafana)? yes / no
- This is a full time, on site role (9 AM to 6 PM). Are you comfortable and available to work on site during these hours? yes / no
- Are you currently based in Egypt? If not, are you willing to relocate to Egypt for this role? choose one
- How would you rate your English proficiency (written and spoken)? written answer
- Describe your experience managing a NOC or application monitoring team. What was the team size and scope of systems monitored? written answer
- Walk us through how you handle incident detection, triage, and escalation. What runbooks or playbooks have you used? written answer
- What observability or monitoring tools have you worked with, and how did you use them to track system availability and performance? written answer
- Tell us about your experience establishing or operating Site Reliability Engineering (SRE) practices. What impact did it have? written answer
- Describe a time you identified a recurring incident pattern and how you addressed the root cause. written answer
- How do you typically produce and use operational health reports or SLO tracking to drive improvements? written answer
- Briefly describe your experience working within a regulated environment (banking, finance, or similar), including any compliance requirements relevant to monitoring operations. written answer
- What is your current monthly net salary? written answer
- What are your expected monthly net salary requirements? written answer
- When can you start? (Notice Period) written answer