Infrastructure Subject Matter Expert for Bcdr and Dr Automation

Open 19d

Role overview:

The infrastructure SME plays a critical role in ensuring that the underlying IT infrastructure fully supports business continuity and disaster recovery (BCDR) objectives, with a strong focus on DR automation. This role bridges infrastructure and automation teams, ensuring resilience, scalability, and seamless failover/failback execution across all infrastructure layers.

Key responsibilities

  1. Infrastructure architecture & readiness
    • Review and validate the end-to-end infrastructure architecture supporting automated DR failover and failback, including:
    o Network, security, compute, storage, virtualization, containers, and data center components
    • Ensure the design supports high availability, resiliency, and recoverability aligned with business requirements.
  2. DR automation integration
    • Act as the primary bridge between infrastructure teams and the DR automation team, ensuring alignment and seamless collaboration.
    • Review and validate automated failover/failback workflows across infrastructure components, including:
    o Network, security, servers, storage, DNS, virtualization platforms, and container environments
    • Collaborate on the development of pre-failover validation scripts to ensure readiness before execution.
  3. Recovery objectives & capacity planning
    • Review and validate infrastructure-level RTOs, ensuring alignment with application and business recovery requirements.
    • Ensure sufficient capacity and performance within DR sites and automation platforms to support:
    o Full failover scenarios
    o Partial or phased failover scenarios
  4. Technical leadership & engagement
    • Lead and actively participate in technical discussions and workshops across:
    o Discovery
    o Validation
    o Tabletop exercises
    • Provide domain expertise and recommendations to ensure robust infrastructure design and DR strategy alignment.
  5. Performance & validation
    • Oversee and validate infrastructure performance testing during and after DR failover/failback activities.
    • Ensure that systems meet defined performance benchmarks and recovery objectives post-recovery.
  6. Compliance & audit readiness
    • Review and ensure adherence to audit and regulatory requirements, particularly around:
    o Logging
    o Monitoring
    o Traceability of DR activities
    • Support audit readiness by ensuring proper documentation and controls are in place.
  7. Cross-functional collaboration
    • Collaborate with application, network, security, database, and business teams to ensure end-to-end alignment.
    • Coordinate with stakeholders to ensure dependencies are properly managed across infrastructure and application layers.
  8. Continuous improvement & optimization
    • Identify opportunities to optimize infrastructure resilience, performance, and cost efficiency.
    • Drive continuous improvement initiatives based on test results, incidents, and evolving business needs.

Requirements

  • Strong expertise in enterprise infrastructure design and operations (network, compute, storage, virtualization, cloud).
  • Hands-on experience with disaster recovery architectures and DR automation tools.
  • Deep understanding of failover/failback mechanisms and infrastructure dependencies.
  • Experience in capacity planning, performance testing, and high availability design.
  • Knowledge of regulatory and compliance requirements related to DR and infrastructure.
  • Strong stakeholder communication and cross-team coordination skills.

Preferred qualifications

  • Experience in large-scale BCDR and DR automation programs.
  • Certifications in infrastructure technologies, cloud platforms, or DR/BCDR frameworks.