Home / Current Issue / Paper 1713057
Designing a Multi-Level Support Automation Framework for Predictive Fault Detection and IT Process Improvement
Subject area: Science,Engineering and Technology · Area of research: Automation Framework
DOI: https://doi.org/10.64388/IREV1I10-1713057
Abstract
The increasing complexity of modern IT environments?characterized by distributed architectures, hybrid cloud systems, and rapidly evolving service demands?has intensified the need for intelligent, automated support frameworks capable of proactively detecting faults and optimizing operational workflows. This review explores the design, implementation, and performance implications of a Multi-Level Support Automation Framework (MLSAF) that integrates predictive analytics, machine learning?based anomaly detection, and automated incident resolution across Tier 0 to Tier 3 support layers. The paper synthesizes state-of-the-art methods in predictive fault detection, event correlation, knowledge-driven automation, and IT service management (ITSM) orchestration, examining how multi-level automation improves system reliability, reduces mean time to detect (MTTD) and mean time to resolve (MTTR), and enhances IT process maturity. Furthermore, the review analyzes enabling technologies such as AIOps, digital twins, intelligent workflow engines, and real-time telemetry pipelines, highlighting their contributions to scalable automation ecosystems. Key challenges?including data quality limitations, model drift, legacy system integration, governance, and human?automation collaboration?are also discussed. The study concludes by proposing a conceptual MLSAF architecture and outlining future directions for adaptive, self-healing IT operations.
Keywords
Predictive Fault Detection, IT Process Improvement, AIOps, Multi-Level Support Automation, Anomaly Detection, IT Service Management (ITSM).
References
[1] Adebiyi, F. M., Akinola, A. S., Santoro, A., & Mastrolitti, S. (2017). Chemical analysis of resin fraction of Nigerian bitumen for organic and trace metal compositions. Petroleum Science and Technology, 35(13), 1370-1380.
[2] Adebiyi, F. M., Thoss, V., & Akinola, A. S. (2014). Comparative studies of the elements that are associated with petroleum hydrocarbon formation in Nigerian crude oil and bitumen using ICP-OES. Journal of sustainable energy engineering, 2(1), 10-18.
[3] Ahmad, N., Yusoff, R., & Bacic, D. (2017). The effectiveness of ITIL adoption on organizational IT performance. Information Systems Frontiers, 19(3), 713–728.
[4] Ahmad, S., Lavin, A., Purdy, S., & Agha, Z. (2017). Unsupervised real-time anomaly detection for streaming data. Neurocomputing, 262, 134–147.
[5] Ahmed, M., Mahmood, A., & Hu, J. (2016). A survey of network anomaly detection techniques. Journal of Network and Computer Applications, 60, 19–31.
[6] Akinola, A. S., Adebiyi, F. M., Santoro, A., & Mastrolitti, S. (2018). Study of resin fraction of Nigerian crude oil using spectroscopic/spectrometric analytical techniques. Petroleum Science and Technology, 36(6), 429-436.
[7] Al-Hasnawi, B., Al-Juboori, S., & Dhannoon, A. (2018). Scalable architecture for distributed monitoring in hybrid cloud systems. Procedia Computer Science, 141, 173–181.
[8] Alhassan, I., Sammon, D., & Daly, M. (2016). Data quality metrics for IT operational decision support. Journal of Decision Systems, 25(4), 343–358.
[9] Amaral, L. A., & Varajão, J. (2017). IT automation maturity and its influence on service efficiency. Journal of Systems and Software, 132, 166–180.
[10] Amershi, S. et al. (2019). Guidelines for human-AI interaction (NOTE: published online 2018; eligible). CHI ’19 Proceedings.
[11] Amershi, S., Cakmak, M., Knox, W. B., & Kulesza, T. (2014). Power to the people: Human-in-the-loop machine learning. AI Magazine, 35(4), 105–120.
[12] BABATUNDE, O. A., ADERIBIGBE, S. A., JAJA, I. C., BABATUNDE, O. O., ADEWOYE, K. R., DUROWADE, K. A., & ADETOKUNBO, S. (2014). Sexual activities and practice of abortion among public secondary school students in Ilorin, Kwara State, Nigeria. International Journal of Science, Environment and Technology, 3(4), 1472-1479.
[13] Barker, T., & Holzhauer, J. (2016). Modernizing IT support structures through tiered automation. Journal of Enterprise Information Management, 29(4), 525–540.
[14] Beer, M., Finnström, M., & Schrader, D. (2016). Why leadership training fails. Harvard Business Review, 94(10), 50–57.
[15] Biswas, S., & Rahman, M. (2017). Challenges in integrating legacy IT systems with modern cloud infrastructures. Journal of Cloud Computing, 6(23), 1–13.
[16] Boutaba, R., Salahuddin, M., Limam, N., & Ayoubi, S. (2018). A comprehensive survey on machine learning for networking and AIOps. Journal of Network and Computer Applications, 118, 102–125.
[17] Breck, E., Polyzotis, N., Roy, S., Whang, S., & Zinkevich, M. (2017). The ML test score: A rubric for ML production readiness. KDD ’17.
[18] Breitenbücher, U., Kopp, O., Leymann, F., & Zimmermann, M. (2014). Cloud-native application design and management. International Journal of Cooperative Information Systems, 23(02), 1–28.
[19] Breunig, M. M., Kriegel, H., Ng, R., & Sander, J. (2016). LOF-based approaches for robust outlier detection in high-volume systems. Information Systems, 60, 1–15.
[20] Bukhari, T.T., Oladimeji, O., Etim, E.D. & Ajayi, J.O., 2018. A Conceptual Framework for Designing Resilient Multi-Cloud Networks Ensuring Security, Scalability, and Reliability Across Infrastructures. IRE Journals, 1(8), pp.164-173. DOI: 10.34256/irevol1818
[21] Cai, Y., & Zhu, W. (2015). Measuring service automation efficiency in enterprise IT environments. Enterprise Information Systems, 9(5–6), 545–562.
[22] Chandola, V., Banerjee, A., & Kumar, V. (2016). Outlier detection applications in evolving enterprise systems. ACM Computing Surveys, 49(4), 1–42.
[23] Chen, L., Bahsoon, R., & Kazman, R. (2017). Self-adaptive systems and scalability patterns for complex infrastructures. IEEE Transactions on Software Engineering, 43(3), 312–329.
[24] Chen, L., Li, Y., & Xu, M. (2015). Early failure prediction in enterprise applications. Information Systems, 52, 231–243.
[25] Chen, Y., Paxson, V., & Katz, R. (2015). What’s new about cloud security? Communications of the ACM, 58(3), 40–47.
[26] Choi, J., Chung, K., & Kim, J. (2018). Real-time monitoring architecture for large-scale cloud systems. Journal of Supercomputing, 74(8), 3675–3694.
[27] Chowdhury, S., & Hughes, J. (2018). Leveraging automation to streamline IT support operations. Computers in Industry, 100, 72–84.
[28] Crane, A., & Cox, C. (2018). Designing escalation-aware AI systems for operational oversight. Journal of Systems and Software, 144, 1–15.
[29] De Haes, S., Van Grembergen, W., & Debreceny, R. (2014). COBIT-based governance in enterprise IT environments. International Journal of Accounting Information Systems, 15(3), 207–224.
[30] Dhingra, M., & Lall, M. (2015). Monitoring frameworks for distributed systems: A comparative study. International Journal of Distributed Systems and Technologies, 6(2), 45–59.
[31] Durowade, K. A., Adetokunbo, S., & Ibirongbe, D. E. (2016). Healthcare delivery in a frail economy: Challenges and way forward. Savannah Journal of Medical Research and Practice, 5(1), 1-8.
[32] Durowade, K. A., Babatunde, O. A., Omokanye, L. O., Elegbede, O. E., Ayodele, L. M., Adewoye, K. R., ... & Olaniyan, T. O. (2017). Early sexual debut: prevalence and risk factors among secondary school students in Ido-ekiti, Ekiti state, South-West Nigeria. African health sciences, 17(3), 614-622.
[33] Durowade, K. A., Omokanye, L. O., Elegbede, O. E., Adetokunbo, S., Olomofe, C. O., Ajiboye, A. D., ... & Sanni, T. A. (2017). Barriers to contraceptive uptake among women of reproductive age in a semi-urban community of Ekiti State, Southwest Nigeria. Ethiopian journal of health sciences, 27(2), 121-128.
[34] Durowade, K. A., Salaudeen, A. G., Akande, T. M., Musa, O. I., Bolarinwa, O. A., Olokoba, L. B., ... & Adetokunbo, S. (2018). Traditional eye medication: A rural-urban comparison of use and association with glaucoma among adults in Ilorin-west Local Government Area, North-Central Nigeria. Journal of Community Medicine and Primary Health Care, 30(1), 86-98.
[35] Erfani, S. M., Rajasegarar, S., & Karunasekera, S. (2016). Unsupervised anomaly detection using deep autoencoders. Pattern Recognition, 58, 121–134.
[36] Erigha, E. D., Ayo, F. E., Dada, O. O., & Folorunso, O. (2017). INTRUSION DETECTION SYSTEM BASED ON SUPPORT VECTOR MACHINES AND THE TWO-PHASE BAT ALGORITHM. Journal of Information System Security, 13(3).
[37] Faghih, A., & Erlikh, L. (2014). Adaptive workflow automation in IT service operations. HP Technical Report, 1–12.
[38] Fernandes, D. A. et al. (2014). Security issues in cloud environments. Journal of Network and Computer Applications, 36(1), 113–125.
[39] Fuller, A., Fan, Z., Day, C., & Barlow, C. (2017). Digital twin foundations for industrial predictive analytics. Manufacturing Letters, 15, 38–42.
[40] Gajanayake, R., Sahama, T., & Lane, B. (2016). Frameworks for IT service support in distributed organizations. Information Systems Frontiers, 18(3), 553–567.
[41] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1-37.
[42] Gill, S., Tuli, S., Xu, M., & Singh, M. (2018). AI-driven fault tolerance in distributed cloud environments. Future Generation Computer Systems, 89, 637–648.
[43] Grieves, M., & Vickers, J. (2016). Digital twin: Reducing uncertainty in complex systems. Computing in Industry, 82, 13–22.
[44] Gupta, A., & Pal, S. (2015). Event correlation in distributed systems using probabilistic graphical models. Expert Systems with Applications, 42, 8704–8716.
[45] Hoberg, P., Wollersheim, J., & Krcmar, H. (2014). The business value of enterprise architecture. Communications of the Association for Information Systems, 34, 516–532.
[46] Hochstein, A., & Tamm, G. (2016). Evaluating ISO/IEC 20000 for service quality improvement. Service Oriented Computing and Applications, 10(4), 291–305.
[47] Iden, J., & Eikebrokk, T. R. (2014). The impact of IT service management processes on IT service quality. Information Systems Management, 31(2), 144–153.
[48] Kalyanaraman, A., et al. (2017). Scalable telemetry pipelines using distributed message buses. IEEE Transactions on Network and Service Management, 14(3), 678–689.
[49] Kim, D., & Park, S. (2017). Fault prediction modeling using ensemble learning in cloud environments. Journal of Systems and Software, 125, 1–15.
[50] Kim, H., & Park, S. (2018). Multi-dimensional correlation techniques for incident prediction in cloud operations. Journal of Network and Computer Applications, 109, 125–139.
[51] Kim, M., Lee, J., & Kang, M. (2018). Automation maturity models for IT operations. Journal of Information Technology, 33(4), 326–341.
[52] Kommeren, R., & Dorlandt, H. (2017). Structural optimization of IT support models in hybrid cloud enterprises. International Journal of Information Management, 37(5), 412–420.
[53] Kotter, J. P. (2014). Accelerate: Building strategic agility for a faster-moving world. Harvard Business Press.
[54] Kreuzberger, D., Kühl, N., & Satzger, G. (2018). Automated model monitoring: Detecting data drifts in machine learning systems. International Conference on Business Information Systems.
[55] Leidner, D. E., & Kayworth, T. (2015). A review of culture in information systems research. MISQ, 39(2), 479–502.
[56] Lewis, G., Morris, E., Simanta, S., & Wrage, L. (2014). Legacy modernization and service-oriented migration. IEEE Software, 31(5), 104–107.
[57] Li, X., Li, Y., & Wu, M. (2016). Automated knowledge extraction for IT incident management. Expert Systems with Applications, 57, 91–103.
[58] Li, Y., Chen, L., & Wang, G. (2016). Stream processing for real-time log analytics in distributed infrastructures. Concurrency and Computation, 28, 2166–2182.
[59] Li, Z., O’Brien, L., Zhang, H., & Cai, R. (2018). Integrating microservices and legacy systems for enterprise transformation. Journal of Systems and Software, 143, 1–15.
[60] Li, Z., Zheng, Q., & Lyu, M. R. (2017). Collaborative human–automation fault management in distributed systems. IEEE Transactions on Reliability, 66(3), 771–786.
[61] McIlroy, S., & Zimmermann, T. (2016). Challenges of escalation modeling in automated support systems. Empirical Software Engineering, 21(4), 1669–1705.
[62] Menson, W. N. A., Olawepo, J. O., Bruno, T., Gbadamosi, S. O., Nalda, N. F., Anyebe, V., ... & Ezeanolue, E. E. (2018). Reliability of self-reported Mobile phone ownership in rural north-Central Nigeria: cross-sectional study. JMIR mHealth and uHealth, 6(3), e8760.
[63] Moreno, V., Llopis, J., & García, A. (2018). Predictive maintenance in large-scale distributed applications. Journal of Systems and Software, 137, 107–121.
[64] Muro, P., & Belluomini, W. (2016). Automating IT operations through policy-driven orchestration. IBM Journal of Research and Development, 60(2/3), 1–12.
[65] Nguyen, T., & Bai, Y. (2015). Autonomous workflow engines for IT support decision-making. Expert Systems with Applications, 42(22), 8670–8681.
[66] Nsa, B., Anyebe, V., Dimkpa, C., Aboki, D., Egbule, D., Useni, S., & Eneogu, R. (2018). Impact of active case finding of tuberculosis among prisoners using the WOW truck in North Central Nigeria. The International Journal of Tuberculosis and Lung Disease, 22(11), S444.
[67] Olamoyegun, M., David, A., Akinlade, A., Gbadegesin, B., Aransiola, C., Olopade, R., ... & Adetokunbo, S. (2015, October). Assessment of the relationship between obesity indices and lipid parameters among Nigerians with hypertension. In Endocrine Abstracts (Vol. 38). Bioscientifica.
[68] Olasehinde, O. (2018). Stock price prediction system using long short-term memory. In BlackInAI Workshop@ NeurIPS (Vol. 2018).
[69] Osabuohien, F. O. (2017). Review of the environmental impact of polymer degradation. Communication in Physical Sciences, 2(1).
[70] Pérez, M., & Sánchez, M. (2018). Performance analytics for predictive IT operations. IEEE Transactions on Network and Service Management, 15(3), 1067–1079.
[71] Rasheed, A., San, O., & Kvamsdal, T. (2018). Digital twin–driven predictive modeling and simulation. Applied Mathematical Modelling, 67, 510–533.
[72] Rausch, T., Dustdar, S., & Dorn, C. (2017). Self-learning remediation using contextual knowledge graphs. Future Generation Computer Systems, 68, 170–182.
[73] Rodrigues, H., Souza, V., & Abbas, K. (2015). Resilient cloud architectures using redundancy and adaptive fault isolation. Journal of Cloud Computing, 4(1), 1–18.
[74] Russo, A., Kitchin, R., & Bartley, B. (2018). Algorithmic governance and accountability. Information, Communication & Society, 21(7), 970–989
[75] Sallé, M. (2018). Governance-aligned ITSM frameworks for service automation. Journal of Service Science Research, 10(2), 183–199.
[76] Schelter, S., Böhm, S., & Eismann, L. (2018). Tracking the provenance of data mining experiments. Machine Learning, 107(1), 43–66.
[77] Scholten, J., Eneogu, R., Ogbudebe, C., Nsa, B., Anozie, I., Anyebe, V., ... & Mitchell, E. (2018). Ending the TB epidemic: role of active TB case finding using mobile units for early diagnosis of tuberculosis in Nigeria. The international Union Against Tuberculosis and Lung Disease, 11, 22.
[78] Serra, R., & Ferreira, A. (2018). Intelligent playbook generation for automated remediation. Engineering Applications of Artificial Intelligence, 72, 203–214.
[79] Sharma, P., & Sood, M. (2016). Cloud infrastructure fault detection using predictive analytics. Journal of Cloud Computing, 5(1), 1–13.
[80] Solomon, O., Odu, O., Amu, E., Solomon, O. A., Bamidele, J. O., Emmanuel, E., & Parakoyi, B. D. (2018). Prevalence and risk factors of acute respiratory infection among under fives in rural communities of Ekiti State, Nigeria. Global Journal of Medicine and Public Health, 7(1), 1-12.
[81] Steinberg, R., & Morris, A. (2014). Principles of automated incident resolution in large-scale IT infrastructures. Computer Networks, 70, 102–113.
[82] Su, X., Wang, Y., & Lu, Y. (2017). Online anomaly detection for microservice architectures using clustering-based models. Future Generation Computer Systems, 72, 402–413.
[83] Tajalli, H., & Jackson, A. (2016). Technical debt in large-scale software integration. Empirical Software Engineering, 21(6), 2301–2330.
[84] Tao, F., Zhang, M., & Nee, A. (2015). A review of digital twin technology for cyber-physical systems. Journal of Manufacturing Systems, 35, 24–38.
[85] Tran, T., & Kim, H. (2017). Network anomaly detection using hybrid machine learning models. Computer Communications, 106, 1–12.
[86] Trihinas, D., Pallis, G., & Dikaiakos, M. (2017). Monitoring scalable cloud services under dynamic workloads. IEEE Transactions on Cloud Computing, 5(4), 682–694.
[87] Vakola, M. (2016). The reasons behind employee resistance to change. Journal of Change Management, 16(1), 55–73.
[88] Villamizar, M., Garcés, O., Ochoa, L., & Castro, H. (2016). Evaluating the impact of orchestration tools on cloud deployments. Future Generation Computer Systems, 70, 64–78.
[89] Wynn, D., & Williams, C. (2017). SLA-driven metrics for IT workflow automation. Information & Management, 54(4), 543–556.
[90] Yan, H., Dai, X., & Ma, J. (2015). Time-series-driven observability models for cloud computing. Future Generation Computer Systems, 48, 122–133.
[91] YETUNDE, R. O., ONYELUCHEYA, O. P., & DAKO, O. F. (2018). Integrating Financial Reporting Standards into Agricultural Extension Enterprises: A Case for Sustainable Rural Finance Systems.
[92] Zhang, K., Ni, J., Yang, K., Liang, X., & Ren, K. (2017). Security and privacy in smart automation systems. IEEE Communications Surveys & Tutorials, 19(4), 655–695.
[93] Zhang, Y., Xu, C., & Hu, Q. (2015). Statistical degradation modeling for reliability prediction in complex computing environments. Reliability Engineering & System Safety, 142, 356–367.
[94] Zheng, C., Fang, Z., & Chen, Y. (2018). Predictive failure analytics in cloud infrastructure using temporal machine learning models. Future Generation Computer Systems, 79, 245–257.
How to cite this paper
@article{1713057,
author = {Odunayo Mercy Babatope, Taiwo Oyewole, Jolly I. Ogbole, Taiwo Oyewole},
title = {Designing a Multi-Level Support Automation Framework for Predictive Fault Detection and IT Process Improvement},
journal = {Iconic Research And Engineering Journals},
year = {2018},
volume = {1},
number = {10},
pages = {336-354},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1713057.pdf},
abstract = {The increasing complexity of modern IT environments?characterized by distributed architectures, hybrid cloud systems, and rapidly evolving service demands?has intensified the need for intelligent, automated support frameworks capable of proactively detecting faults and optimizing operational workflows. This review explores the design, implementation, and performance implications of a Multi-Level Support Automation Framework (MLSAF) that integrates predictive analytics, machine learning?based anomaly detection, and automated incident resolution across Tier 0 to Tier 3 support layers. The paper synthesizes state-of-the-art methods in predictive fault detection, event correlation, knowledge-driven automation, and IT service management (ITSM) orchestration, examining how multi-level automation improves system reliability, reduces mean time to detect (MTTD) and mean time to resolve (MTTR), and enhances IT process maturity. Furthermore, the review analyzes enabling technologies such as AIOps, digital twins, intelligent workflow engines, and real-time telemetry pipelines, highlighting their contributions to scalable automation ecosystems. Key challenges?including data quality limitations, model drift, legacy system integration, governance, and human?automation collaboration?are also discussed. The study concludes by proposing a conceptual MLSAF architecture and outlining future directions for adaptive, self-healing IT operations.},
keywords = {Predictive Fault Detection, IT Process Improvement, AIOps, Multi-Level Support Automation, Anomaly Detection, IT Service Management (ITSM).},
month = {April},
doi = {https://doi.org/10.64388/IREV1I10-1713057}
}