Senior Disaster Recovery Analyst
Summary
Designs and executes real-world disaster recovery tests for cloud-based trading systems, using AI tools to automate evidence gathering and improve resilience against outages.
- AI-driven alert correlation and remediation guidance running today, not on a roadmap
- Automated evidence-gathering workflows replacing manual audit prep across the business
- A culture that treats "we tested it and it broke" as useful information, not a failure
- Design and execute DR exercises for business-critical applications — database failures, regional outages, SaaS provider disruptions, live failover — and report honestly on what held up and what didn't
- Work hands-on with DevOps, WinOps, SRE, engineering, product, operations, and SaaS owners to test recovery plans against how systems actually behave under stress
- Feed every incident and near-miss straight back into the DR plan, so the same failure never catches us twice.
- Conduct business impact analyses that pressure-test RTO and RPO targets against real operational data, not last year's assumptions
- Push back when a recovery target looks good on paper but wouldn't survive contact with a real outage
- Maintain and audit DR strategies, runbooks, test records, BIAs, and recovery evidence so they're accurate today, not just accurate when they were written
- Build dashboards that give leadership a real answer to "are we actually ready?" — not a static slide from last quarter
- Model failure scenarios, generate exercise briefs, and simulate impact using AI tooling instead of building everything from a blank template
- Automate evidence gathering and gap surfacing so your time goes to the failures that need judgment, not paperwork
- Analyse historical incidents for patterns humans tend to miss under deadline pressure
- Flag it when a new system goes into production without a tested recovery plan — even when nobody asked you to check
- Translate what's actually at risk into language that works for an infrastructure engineer and for a C-suite leader, in the same conversation if you have to
- You handle routine DR gaps and exercise findings without waiting to be told what to do next — and when you figure something out, you make sure the rest of the team doesn't have to learn it the hard way too.
- Tabletop discussions are a start, not a finish line. You've run real failover tests where something actually broke, and you know the difference between a plan that reads well and a plan that survives contact with reality.
- You reach for AI tooling to model scenarios, draft exercise briefs, and chase down evidence gaps — not because it's expected of you, but because it's obviously the faster, better way to work.
- Your runbooks, BIAs, and recovery narratives are precise enough that someone else could execute them without you in the room, and current enough that they'd actually work if they had to.
- You can tell an engineer exactly what broke and tell a director exactly what it means for the business — and you know which one the room in front of you needs.
- 4+ years in disaster recovery, business continuity, infrastructure resilience, or cloud operations
- Hands-on experience with AWS, GCP, Azure, or equivalent, including cloud-native failure modes
- Practical experience running real DR exercises or failover tests — not only writing plans or facilitating tabletops
- Experience conducting or contributing to business impact analyses, especially in cloud or hybrid environments
- Working knowledge of ITIL and business continuity frameworks, applied in practice
- Comfort using AI tools for analysis, reporting, simulation, and evidence automation
- A relevant certification held or actively in progress — CBCP, MBCI, AWS/GCP Professional, or equivalent
- Preparing DR evidence for Compliance, Risk, or external audit reviews
- Building DR readiness dashboards or AI-assisted resilience reporting
- Integrating AI-driven monitoring or anomaly detection into infrastructure workflows