freehire launches on Product Hunt on 26 August.

Follow →

Cloud Reliability & Recovery Engineer

Open 34d

You will design, implement, and continuously improve Business Continuity Planning (BCP) and Disaster Recovery (DR) capabilities across AWS cloud environments. In this hands-on technical role, you'll build highly available, fault-tolerant, and resilient cloud architecture using Kubernetes for container orchestration and Terraform for infrastructure as code. You'll implement CI/CD pipelines to enable rapid, reliable deployments, and you'll design multi-region failover, backup and restore automation, and recovery testing aligned with industry BCP/DR standards. You'll work closely with security, infrastructure, and application teams to ensure systems can withstand and rapidly recover from disruption, reporting to the Director of Event Response.

Responsibilities

  • Design and implement multi-region, multi-AZ AWS architectures that meet RTO/RPO targets
  • Engineer active-active and active-passive failover patterns using Route 53, Global Accelerator, and CloudFront
  • Build automated DR runbooks and playbooks using AWS Systems Manager Automation and Step Functions
  • Implement chaos engineering practices using AWS Fault Injection Simulator to validate resiliency
  • Architect cross-region replication strategies for S3, DynamoDB Global Tables, RDS, and Aurora Global
  • Review containerized workloads using Kubernetes, ensuring resilience through self-healing, auto-scaling, and multi-cluster or multi-region deployments
  • Administer AWS Backup across all services with policy-based automation
  • Design immutable backup vaults and cross-account/cross-region backup replication pipelines
  • Develop and automate data recovery testing procedures, ensuring integrity and meeting defined SLAs
  • Implement point-in-time recovery for databases and storage; validate via regular restore drills
  • Maintain Business Continuity Plans and Disaster Recovery strategies, tracking RTO and RPO
  • Author and maintain Terraform/CloudFormation templates for all BCP/DR infrastructure components
  • Automate DR testing pipelines through CI/CD tools such as CodePipeline, CodeBuild, and GitHub Actions
  • Write Python/Bash/PowerShell scripts to orchestrate failover, failback, and health-check workflows
  • Manage infrastructure state in AWS Control Tower and implement Landing Zone DR patterns
  • Build CloudWatch dashboards, alarms, and composite alarms for availability and DR-readiness indicators
  • Integrate AWS Health and Personal Health Dashboard events into PagerDuty/OpsGenie alerting workflows
  • Participate in on-call rotations and lead DR incident response; conduct post-incident reviews
  • Develop and maintain runbooks for AWS service degradations, regional outages, and data corruption events
  • Conduct regular BCP/DR tabletop exercises and full failover simulations to validate recovery procedures
  • Ensure DR controls meet SOC 2, ISO 22301, NIST 800-53, and HIPAA/PCI requirements as applicable
  • Maintain current and accurate DR documentation including BIAs, BCPs, DRP runbooks, and recovery evidence
  • Collaborate with audit and compliance teams to provide DR evidence and remediation tracking

Requirements

  • 5+ years in cloud infrastructure, SRE, or IT disaster recovery engineering roles
  • 3+ years of hands-on AWS experience in production environments at scale
  • Proven delivery of multi-region DR architectures with defined and tested RTO/RPO targets
  • Expert-level proficiency with core AWS resilience services
  • Strong scripting skills: Python, Bash, or PowerShell for automation and orchestration
  • Experience with Infrastructure as Code: Terraform and/or AWS CloudFormation
  • Solid understanding of networking fundamentals: VPC, TGW, Direct Connect, VPN, DNS failover
  • Excellent written and verbal communication; able to produce executive-level DR reports
  • AWS Certified Solutions Architect – Professional or AWS Certified DevOps Engineer – Professional preferred
  • AWS Certified Advanced Networking – Specialty certification preferred
  • Experience with AWS Resilience Hub for automated resilience assessments and policy enforcement preferred
  • Familiarity with CloudEndure / AWS Elastic Disaster Recovery for workload replication preferred
  • Knowledge of Kubernetes-based DR such as EKS multi-region, Velero backups, and ArgoCD GitOps failover preferred
  • Hands-on experience with serverless DR patterns such as Lambda, API Gateway, and DynamoDB preferred

See also

Tailor your CV for this role?

We couldn't check your fit for this role — add a CV to your profile to see it next time.

A new version of freehire is available