Home / Current Issue / Paper 1718609
AI-Powered Disaster Recovery Planning and Failover Optimization for Mission-Critical Enterprise Systems
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence
DOI: 10.64388/IREV9I12-1718609
Abstract
Mission-critical enterprise systems now operate through hybrid combinations of private data centres, public cloud zones, edge services, identity platforms, databases, application programming interfaces, and third-party digital supply chains. Conventional disaster recovery planning, while still necessary, is increasingly insufficient because static runbooks cannot anticipate fast-moving failures, cascading dependencies, ransomware disruptions, configuration drift, and volatile workload peaks. This review examines how artificial intelligence can strengthen disaster recovery planning and failover optimization through predictive failure analysis, anomaly detection, automated root cause reasoning, workload forecasting, policy-aware orchestration, and human-supervised remediation. A structured narrative review method was applied to recent research and standards published between 2020 and 2025, with emphasis on AIOps, cloud reliability, critical infrastructure protection, generative artificial intelligence, autoscaling, microservice observability, cyber resilience, and governance. The paper proposes an integrated AI-driven disaster recovery lifecycle that links telemetry ingestion, risk scoring, scenario modelling, failover decisioning, validation, and post-incident learning. The review indicates that the strongest value of AI lies not in replacing continuity professionals but in compressing detection-to-decision time, reducing alert noise, identifying probable blast radius, and recommending recovery actions aligned with recovery time objectives, recovery point objectives, security controls, and service-level commitments. However, the literature also reveals persistent limitations, including explainability gaps, poor data quality, adversarial model risk, over-automation, weak integration with legacy platforms, and uncertain accountability during autonomous failover. The study concludes that AI-powered disaster recovery should be designed as a governed socio-technical capability, combining machine intelligence, resilient architecture, audited automation, and expert approval for high-impact actions.
Keywords
Artificial Intelligence, Disaster Recovery Failover Optimization AIOps, Mission-Critical Systems, Cyber Resilience, Business Continuity.
References
[1] Yigit Y, Ferrag MA, Ghanem MC, Sarker IH, Maglaras LA, Chrysoulas C, Moradpoor N, Tihanyi N, Janicke H (2025) Generative AI and LLMs for critical infrastructure protection: evaluation benchmarks, agentic AI, challenges, and opportunities. Sensors 25(6):1666. https://doi.org/10.3390/s25061666
[2] Ferrag MA, Alwahedi F, Battah A, Cherif B, Mechri A, Tihanyi N, Debbah M, Lestable T (2025) Generative AI in cybersecurity: a comprehensive review of LLM applications and vulnerabilities. Internet of Things and Cyber-Physical Systems 5:1-46. https://doi.org/10.1016/j.iotcps.2025.01.001
[3] Uddin M, Irshad MS, Kandhro IA, Alanazi F, Ahmed F, Maaz M, Hussain S, Ullah SS (2025) Generative AI revolution in cybersecurity: a comprehensive review of threat intelligence and operations. Artificial Intelligence Review 58:236. https://doi.org/10.1007/s10462-025-11219-5
[4] Dawood M, Tu S, Xiao C, Alasmary H, Waqas M, Rehman SU (2023) Cyberattacks and security of cloud computing: a complete guideline. Symmetry 15(11):1981. https://doi.org/10.3390/sym15111981
[5] Soldani J, Brogi A (2022) Anomaly detection and failure root cause analysis in (micro)service-based cloud applications: a survey. ACM Computing Surveys 55(3):1-39. https://doi.org/10.1145/3501297
[6] Chen Y, Xie H, Ma M, Kang Y, Gao X, Shi L, Cao Y, Gao X, Fan H, Wen M, Zeng J, Ghosh S, Zhang X, Zhang C, Lin Q, Rajmohan S, Zhang D, Xu T (2023) Automatic root cause analysis via large language models for cloud incidents. arXiv:2305.15778.
[7] Saha A, Hoi SCH (2022) Mining root cause knowledge from cloud service incident investigations for AIOps. arXiv:2204.11598.
[8] Zhang Y, Guan Z, Qian H, Xu L, Liu H, Wen Q, Sun L, Jiang J, Fan L, Ke M (2021) CloudRCA: a root cause analysis framework for cloud computing platforms. arXiv:2111.03753.
[9] Xu J, Xu Z, Shi B (2022) Deep reinforcement learning based resource allocation strategy in cloud-edge computing system. Frontiers in Bioengineering and Biotechnology 10:908056. https://doi.org/10.3389/fbioe.2022.908056
[10] Gari Y, Monge DA, Pacini E, Mateos C, Garino CG (2021) Reinforcement learning-based application autoscaling in the cloud: a survey. Engineering Applications of Artificial Intelligence 102:104288. https://doi.org/10.1016/j.engappai.2021.104288
[11] Xue S, Qu C, Shi X, Liao C, Zhu S, Tan X, Ma L, Wang S, Wang S, Hu Y, Lei L, Zheng Y, Li J, Zhang J (2022) A meta reinforcement learning approach for predictive autoscaling in the cloud. arXiv:2205.15795.
[12] Fettes Q, Karanth A, Bunescu R, Beckwith B, Subramoney S (2023) Reclaimer: a reinforcement learning approach to dynamic resource allocation for cloud microservices. arXiv:2304.07941.
[13] Abdel Khaleq A, Ra I (2023) Intelligent microservices autoscaling module using reinforcement learning. Cluster Computing 26:2789-2800. https://doi.org/10.1007/s10586-023-03999-8
[14] Nobre J, Pires EJS, Reis A (2023) Anomaly detection in microservice-based systems. Applied Sciences 13(13):7891. https://doi.org/10.3390/app13137891
[15] Kohyarnejadfard I, Aloise D, Azhari SV, Dagenais MR (2022) Anomaly detection in microservice environments using distributed tracing data analysis and NLP. Journal of Cloud Computing 11:25. https://doi.org/10.1186/s13677-022-00296-4
[16] Notaro P, Cardoso J, Gerndt M (2020) A systematic mapping study in AIOps. arXiv:2012.09108.
[17] Jiang Z, Li T, Zhang Z, Ge J, You J, Li L (2021) A survey on log research of AIOps: methods and trends. Mobile Networks and Applications 26:2353-2364. https://doi.org/10.1007/s11036-021-01832-3
[18] Zolanvari M, Ghubaish A, Jain R (2022) ADDAI: anomaly detection using distributed AI. arXiv:2205.01231.
[19] Fernando D, Rodriguez MA, Buyya R (2024) iAnomaly: a toolkit for generating performance anomaly datasets in edge-cloud integrated computing environments. arXiv:2411.02868.
[20] National Institute of Standards and Technology (2023) Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1. https://doi.org/10.6028/NIST.AI.100-1
[21] National Institute of Standards and Technology (2024) The NIST Cybersecurity Framework 2.0. National Institute of Standards and Technology, Gaithersburg.
[22] European Union Agency for Cybersecurity (2024) ENISA Threat Landscape 2024. ENISA, Athens.
[23] OWASP Foundation (2025) OWASP Top 10 for Large Language Model Applications 2025. OWASP, online publication.
[24] International Organization for Standardization and International Electrotechnical Commission (2022) ISO/IEC 27001:2022 Information security, cybersecurity and privacy protection - Information security management systems - Requirements. ISO, Geneva.
[25] International Organization for Standardization and International Electrotechnical Commission (2023) ISO/IEC 42001:2023 Information technology - Artificial intelligence - Management system. ISO, Geneva.
[26] Amazon Web Services (2024) AWS Well-Architected Framework: Reliability Pillar. Amazon Web Services, Seattle.
[27] Microsoft (2024) Azure Well-Architected Framework: Reliability. Microsoft, Redmond.
[28] Google Cloud (2023) Disaster recovery planning guide. Google Cloud Architecture Center, Mountain View.
[29] Cybersecurity and Infrastructure Security Agency (2023) Cross-Sector Cybersecurity Performance Goals. CISA, Washington, DC.
[30] International Organization for Standardization (2022) ISO 22361:2022 Security and resilience - Crisis management - Guidelines. ISO, Geneva.
How to cite this paper
@article{1718609,
author = {Tahseen Zafar},
title = {AI-Powered Disaster Recovery Planning and Failover Optimization for Mission-Critical Enterprise Systems},
journal = {Iconic Research And Engineering Journals},
year = {2026},
volume = {9},
number = {12},
pages = {581-591},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1718609.pdf},
abstract = {Mission-critical enterprise systems now operate through hybrid combinations of private data centres, public cloud zones, edge services, identity platforms, databases, application programming interfaces, and third-party digital supply chains. Conventional disaster recovery planning, while still necessary, is increasingly insufficient because static runbooks cannot anticipate fast-moving failures, cascading dependencies, ransomware disruptions, configuration drift, and volatile workload peaks. This review examines how artificial intelligence can strengthen disaster recovery planning and failover optimization through predictive failure analysis, anomaly detection, automated root cause reasoning, workload forecasting, policy-aware orchestration, and human-supervised remediation. A structured narrative review method was applied to recent research and standards published between 2020 and 2025, with emphasis on AIOps, cloud reliability, critical infrastructure protection, generative artificial intelligence, autoscaling, microservice observability, cyber resilience, and governance. The paper proposes an integrated AI-driven disaster recovery lifecycle that links telemetry ingestion, risk scoring, scenario modelling, failover decisioning, validation, and post-incident learning. The review indicates that the strongest value of AI lies not in replacing continuity professionals but in compressing detection-to-decision time, reducing alert noise, identifying probable blast radius, and recommending recovery actions aligned with recovery time objectives, recovery point objectives, security controls, and service-level commitments. However, the literature also reveals persistent limitations, including explainability gaps, poor data quality, adversarial model risk, over-automation, weak integration with legacy platforms, and uncertain accountability during autonomous failover. The study concludes that AI-powered disaster recovery should be designed as a governed socio-technical capability, combining machine intelligence, resilient architecture, audited automation, and expert approval for high-impact actions.},
keywords = {Artificial Intelligence, Disaster Recovery Failover Optimization AIOps, Mission-Critical Systems, Cyber Resilience, Business Continuity.},
month = {June},
doi = {https://doi.org/10.64388/IREV9I12-1718609}
}