Point your AI agent at freehire and let it find you a job.

Get the CLI →

Incident Manager

NewBe an early applicant

About the Role

We are recruiting an Incident Manager based in the UK. The Incident Manager will work 8-hour shifts across a five-day week and provide on-call assistance as part of a four-person rota that delivers 24/7 escalation coverage. Shift schedules may be fixed or rotated, depending on engineers' preferences.

These roles do not require night shifts.

Benefits

  • 33 days annual leave plus your birthday day

  • Salary sacrifice company pension scheme

  • Personal life insurance (3x your salary) and income protection

  • Health insurance with options to add your family

  • Dedicated professional development training budget

  • Enhanced company sick pay

  • Enhanced family leave policy

  • 2 paid volunteer days per year

  • Electric Car Scheme (salary sacrifice)

Salary Range: £51,000 - £70,000

As a Digital Incident Manager, you will:

  • Assume the Incident Commander role for all P1 and P2 incidents: own the bridge call, drive resolution, and coordinate cross-functional resolver groups.

  • Establish roles on incident bridges (scribe, technical lead, communications lead) and enforce time-boxed troubleshooting with 30-minute checkpoints to avoid stagnation.

  • Make escalation decisions: page on-call engineers, engage leadership or vendors, notify compliance, or execute rollback/failover/service-degradation procedures.

  • Send initial stakeholder notifications within the SLA and maintain a consistent cadence: technical details for engineering; business impact for leadership.

  • Coordinate with Compliance and Regulatory teams for impact notifications and initiate customer-facing communications (app banners, status page updates) when required.

  • Enforce monitoring coverage requirements for all Caesars Digital production services, ensure no service goes to production without adequate monitoring and alerting in place.

  • Continuously tune alert thresholds based on feedback from Digital System Support Engineers, reducing false positives and alert fatigue while driving the alert signal-to-noise ratio above 80% actionable.

  • Implement alert deduplication, correlation, and suppression rules to ensure Engineers receive clean, actionable signals rather than noise that degrades response effectiveness.

  • Define and maintain alert severity standards that clearly distinguish P1 vs. P2 vs. P3 vs. informational alerts, ensuring consistent classification across all services.

  • Review "missed detection" findings from the Major Incident Manager's Post-Incident Reviews and build new monitoring coverage to prevent recurrence of undetected issues.

  • Ensure every alert in the ecosystem links to a documented runbook with clear response procedures that Engineers can execute independently.

  • Build and own end-to-end customer journey monitoring covering the critical user flows: Registration → Deposit → Bet Placement → Bet Settlement → Withdrawal, with defined thresholds for success rates and drop-off alerts.

  • Design and implement real-time revenue monitoring dashboards tracking deposit/withdrawal volumes, payment gateway health, and transaction success rates with anomaly detection against expected baselines.

  • Build revenue impact calculation models for use during major incidents, enabling the team to quantify business impact in dollar terms.

  • Define and maintain business KPI dashboards monitoring operational metrics including handle, active users, concurrent sessions, bet volume per minute, and Caesars Rewards pipeline health.

  • Create executive-visible business health dashboards that provide real-time situational awareness during high-revenue events and peak traffic periods.

  • Conduct monthly monitoring audits to assess coverage completeness, alert quality, and identify stale or orphaned alerts for decommissioned services.

  • Build synthetic monitoring for critical customer journeys to validate service availability and performance from the customer's perspective.

  • Collaborate with the Problem Manager to support event readiness by building enhanced dashboards, lowering detection thresholds, and adding event-specific synthetic monitors per the readiness plan.

  • Report on noisy alert sources monthly and drive engineering teams to fix the root causes generating non-actionable alerts.

  • Partner with product and engineering teams to agree on monitoring thresholds and ensure observability is built into the development lifecycle.

EDUCATION AND EXPERIENCE:

  • 5+ years of experience in IT operations, site reliability engineering, observability engineering, or a senior technical monitoring role.

  • Deep hands-on expertise with enterprise monitoring and observability platforms such as Zabbix, Splunk, DynaTrace, New Relic, Datadog, Grafana, or similar tools.

  • Proven experience designing and implementing monitoring strategies for large-scale, distributed, customer-facing digital platforms.

  • Strong understanding of alert engineering principles including threshold tuning, correlation rules, suppression logic, and noise reduction techniques.

  • Experience building business-level monitoring including revenue dashboards, customer journey tracking, and transaction anomaly detection.

  • Ability to translate technical metrics into business impact language for executive stakeholders.

  • Strong analytical skills with experience in defining KPIs, setting baselines, and measuring Mean Time to Detect (MTTD) improvements.

  • Experience working with engineering teams to embed observability into CI/CD pipelines and service delivery processes.

  • Understanding of cloud-native architectures, microservices, containerization, and modern infrastructure patterns.

  • Previous experience in gaming, sports betting, or high-transaction digital commerce environments is strongly preferred.

  • Familiarity with ITIL processes, particularly Event Management, and experience with on-call escalation models.

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience.

See also

Management jobs by country — openings, pay and top skills →

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available