International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1710020

1710020 Vol 2 · Issue 11 Download Paper

Enhancing Enterprise Software Reliability Using Retry Queues and Message Persistence in Event-Driven Cloud Environments

Eseoghene Daniel Erigha Ehimah Obuse Babawale Patrick Okare Abel Chukwuemeke Uzoka Samuel Owoade Noah Ayanbode

Subject area: Science,Engineering and Technology  ·  Area of research: Software Reliability

Abstract

In an era where enterprise software systems are increasingly deployed on cloud platforms and built upon event-driven architectures, ensuring consistent reliability across distributed components becomes a critical concern. These modern architectures promote scalability and responsiveness through asynchronous communication, but they also introduce new complexities in handling transient failures, message delivery guarantees, and fault tolerance. This explores the role of retry queues and message persistence as foundational mechanisms for enhancing software reliability in such environments. Retry queues enable services to automatically attempt message processing again after initial failures, using configurable strategies such as exponential backoff, jitter, and maximum retry limits. These mechanisms help prevent message loss, reduce system downtime, and improve end-to-end transaction success rates. When integrated with dead-letter queues and observability tools, retry queues offer not only recovery but also insight into persistent system weaknesses and transient bottlenecks. Message persistence further strengthens reliability by ensuring that messages are durably stored?often across distributed logs or message brokers?until they are successfully processed or safely discarded. Leveraging technologies such as Apache Kafka, AWS SQS with Dead-Letter Queues, and Azure Service Bus, developers can implement various delivery semantics (at-least-once, exactly-once, at-most-once) suited to different application requirements. Persistence protects against system crashes, network partitions, and service restarts, thereby maintaining data integrity and continuity across the system. This synthesizes architectural best practices, cloud-native tooling, and design patterns for implementing retry logic and persistent messaging in microservice-based systems. It also highlights real-world use cases?including transactional processing, notification systems, and event sourcing?demonstrating how these reliability mechanisms can be effectively employed. Finally, the discussion explores future directions such as AI-assisted retry strategies, serverless queue orchestration, and cross-cloud persistence standards. In conclusion, retry queues and message persistence are indispensable tools for building fault-tolerant, enterprise-grade, event-driven software in dynamic cloud environments.

Keywords

Enterprise, Software reliability, Retry queues, Message persistence, Event-driven, Cloud environments

References

[1] Ajonbadi Adeniyi, H., AboabaMojeed-Sanni, B. and Otokiti, B.O., 2015. Sustaining competitive advantage in medium-sized enterprises (MEs) through employee social interaction and helping behaviours. Journal of Small Business and Entrepreneurship, 3(2), pp.1-16.

[2] Ajonbadi, H.A., Lawal, A.A., Badmus, D.A. and Otokiti, B.O., 2014. Financial control and organisational performance of the Nigerian small and medium enterprises (SMEs): A catalyst for economic growth. American Journal of Business, Economics and Management, 2(2), pp.135-143.

[3] Ajonbadi, H.A., Otokiti, B.O. and Adebayo, P., 2016. The efficacy of planning on organisational performance in the Nigeria SMEs. European Journal of Business and Management, 24(3), pp.25-47.

[4] Akinbola, O.A. and Otokiti, B.O., 2012. Effects of lease options as a source of finance on profitability performance of small and medium enterprises (SMEs) in Lagos State, Nigeria. International Journal of Economic Development Research and Investment, 3(3), pp.70-76.

[5] Amos, A.O., Adeniyi, A.O. and Oluwatosin, O.B., 2014. Market based capabilities and results: inference for telecommunication service businesses in Nigeria. European Scientific Journal, 10(7).

[6] Awe, E.T. and Akpan, U.U., 2017. Cytological study of Allium cepa and Allium sativum.

[7] Awe, E.T., 2017. Hybridization of snout mouth deformed and normal mouth African catfish Clarias gariepinus. Animal Research International, 14(3), pp.2804-2808.

[8] Baek, H., Srivastava, A. and Van der Merwe, J., 2017, May. Cloudsight: A tenant-oriented transparency framework for cross-layer cloud troubleshooting. In 2017 17th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID) (pp. 268-273). IEEE.

[9] Bahill, A.T. and Madni, A.M., 2017. Tradeoff decisions in system design (pp. p-476). Cham: Springer International Publishing.

[10] Beyer, B., Murphy, N.R., Rensin, D.K., Kawahara, K. and Thorne, S., 2018. The site reliability workbook: practical ways to implement SRE. " O'Reilly Media, Inc.".

[11] Blair, G., 2018, July. Complex distributed systems: The need for fresh perspectives. In 2018 IEEE 38th International Conference on Distributed Computing Systems (ICDCS) (pp. 1410-1421). IEEE.

[12] Cardin, C., 2016. Design of a horizontally scalable backend application for online games.

[13] Debski, A., Szczepanik, B., Malawski, M., Spahr, S. and Muthig, D., 2017. A scalable, reactive architecture for cloud applications. IEEE Software, 35(2), pp.62-71.

[14] Dobbelaere, P. and Esmaili, K.S., 2017, June. Kafka versus RabbitMQ: A comparative study of two industry reference publish/subscribe implementations: Industry Paper. In Proceedings of the 11th ACM international conference on distributed and event-based systems (pp. 227-238).

[15] Erik, S. and Emma, L., 2018. Real-Time Analytics with Event-Driven Architectures: Powering Next-Gen Business Intelligence. International Journal of Trend in Scientific Research and Development, 2(4), pp.3097-3111.

[16] Evans-Uzosike, I.O. & Okatta, C.G., 2019. Strategic Human Resource Management: Trends, Theories, and Practical Implications. Iconic Research and Engineering Journals, 3(4), pp.264-270.

[17] Gallipeau, D. and Kudrle, S., 2018. Microservices: Building blocks to new workflows and virtualization. SMPTE Motion Imaging Journal, 127(4), pp.21-31.

[18] Ganesan, A., Alagappan, R., Arpaci-Dusseau, A.C. and Arpaci-Dusseau, R.H., 2017. Redundancy does not imply fault tolerance: Analysis of distributed storage reactions to file-system faults. ACM Transactions on Storage (TOS), 13(3), pp.1-33.

[19] Garrison, J. and Nova, K., 2017. Cloud native infrastructure: patterns for scalable infrastructure and applications in a dynamic environment. " O'Reilly Media, Inc.".

[20] Gunawi, H.S., Suminto, R.O., Sears, R., Golliher, C., Sundararaman, S., Lin, X., Emami, T., Sheng, W., Bidokhti, N., McCaffrey, C. and Srinivasan, D., 2018. Fail-slow at scale: Evidence of hardware performance faults in large production systems. ACM Transactions on Storage (TOS), 14(3), pp.1-26.

[21] Gupta, N., Prakash, A. and Tripathi, R., 2017. Adaptive beaconing in mobility aware clustering based MAC protocol for safety message dissemination in VANET. Wireless Communications and Mobile Computing, 2017(1), p.1246172.

[22] Hukerikar, S. and Engelmann, C., 2017. Resilience design patterns: A structured approach to resilience at extreme scale. arXiv preprint arXiv:1708.07422.

[23] Hussain, F., Anpalagan, A. and Vannithamby, R., 2017. Medium access control techniques in M2M communication: survey and critical review. Transactions on Emerging Telecommunications Technologies, 28(1), p.e2869.

[24] Ibitoye, B.A., AbdulWahab, R. and Mustapha, S.D., 2017. Estimation of drivers’ critical gap acceptance and follow-up time at four–legged unsignalized intersection. CARD International Journal of Science and Advanced Innovative Research, 1(1), pp.98-107.

[25] Jha, S., Formicola, V., Di Martino, C., Dalton, M., Kramer, W.T., Kalbarczyk, Z. and Iyer, R.K., 2017. Resiliency of hpc interconnects: A case study of interconnect failures and recovery in blue waters. IEEE Transactions on Dependable and Secure Computing, 15(6), pp.915-930.

[26] John, V. and Liu, X., 2017. A survey of distributed message broker queues. arXiv preprint arXiv:1704.00411.

[27] Jose, J., 2018. Internet of things. Khanna Publishing House.

[28] Joshi, A., Nagarajan, V., Cintra, M. and Viglas, S., 2018, June. Dhtm: Durable hardware transactional memory. In 2018 ACM/IEEE 45th Annual International Symposium on Computer Architecture (ISCA) (pp. 452-465). IEEE.

[29] Kaur, K., Sharma, D.S. and Kahlon, D.K.S., 2017. Interoperability and portability approaches in inter-connected clouds: A review. ACM Computing Surveys (CSUR), 50(4), pp.1-40.

[30] Kristić, A., Ožegović, J. and Kedžo, I., 2018. Design and Modeling of Self‐Adapting MAC (SaMAC) Protocol with Inconstant Contention Loss Probabilities. Wireless communications and mobile computing, 2018(1), p.6375317.

[31] Laboy, M. and Fannon, D., 2016. Resilience theory and praxis: a critical framework for architecture. Enquiry The ARCC Journal for Architectural Research, 13(1).

[32] Lawal, A.A., Ajonbadi, H.A. and Otokiti, B.O., 2014. Leadership and organisational performance in the Nigeria small and medium enterprises (SMEs). American Journal of Business, Economics and Management, 2(5), p.121.

[33] Lawal, A.A., Ajonbadi, H.A. and Otokiti, B.O., 2014. Strategic importance of the Nigerian small and medium enterprises (SMES): Myth or reality. American Journal of Business, Economics and Management, 2(4), pp.94-104.

[34] Leitao, P., Karnouskos, S., Ribeiro, L., Lee, J., Strasser, T. and Colombo, A.W., 2016. Smart agents in industrial cyber–physical systems. Proceedings of the IEEE, 104(5), pp.1086-1101.

[35] Liu, J., Shen, H. and Narman, H.S., 2018. Popularity-aware multi-failure resilient and cost-effective replication for high data durability in cloud storage. IEEE Transactions on Parallel and Distributed Systems, 30(10), pp.2355-2369.

[36] Marcu, O.C., Costan, A., Antoniu, G., Pérez-Hernández, M.S., Tudoran, R., Bortoli, S. and Nicolae, B., 2017, December. Towards a unified storage and ingestion architecture for stream processing. In 2017 IEEE International Conference on Big Data (Big Data) (pp. 2402-2407). IEEE.

[37] Morar, M., Kumar, A., Abbott, M., Gautam, G.K., Corbould, J. and Bhambhani, A., 2017. Robust Cloud Integration with Azure. Packt Publishing Ltd.

[38] Mukwevho, M.A. and Celik, T., 2018. Toward a smart cloud: A review of fault-tolerance methods in cloud systems. IEEE Transactions on Services Computing, 14(2), pp.589-605.

[39] Narkhede, N., Shapira, G. and Palino, T., 2017. Kafka: the definitive guide: real-time data and stream processing at scale. " O'Reilly Media, Inc.".

[40] Nurkiewicz, T. and Christensen, B., 2016. Reactive programming with RxJava: creating asynchronous, event-based applications. " O'Reilly Media, Inc.".

[41] Nwaimo, C.S., Oluoha, O.M. & Oyedokun, O., 2019. Big Data Analytics: Technologies, Applications, and Future Prospects. Iconic Research and Engineering Journals, 2(11), pp.411-419.

[42] Oberhauser, R. and Stigler, S., 2017. Microflows: enabling agile business process modeling to orchestrate semantically-annotated microservices. In Seventh International Symposium on Business Modeling and Software Design (BMSD 2017), Volume 1 (pp. 19-28).

[43] Ogundipe, F., Sampson, E., Bakare, O.I., Oketola, O. and Folorunso, A., 2019. Digital Transformation and its Role in Advancing the Sustainable Development Goals (SDGs). transformation, 19, p.48.

[44] Oni, O., Adeshina, Y.T., Iloeje, K.F. and Olatunji, O.O., ARTIFICIAL INTELLIGENCE MODEL FAIRNESS AUDITOR FOR LOAN SYSTEMS. Journal ID, 8993, p.1162.

[45] Otokiti, B.O. and Akinbola, O.A., 2013. Effects of lease options on the organizational growth of small and medium enterprise (SME’s) in Lagos State, Nigeria. Asian Journal of Business and Management Sciences, 3(4), pp.1-12.

[46] Otokiti, B.O., 2012. Mode of entry of multinational corporation and their performance in the Nigeria market (Doctoral dissertation, Covenant University).

[47] Otokiti, B.O., 2017. A study of management practices and organisational performance of selected MNCs in emerging market-A Case of Nigeria. International Journal of Business and Management Invention, 6(6), pp.1-7.

[48] Otokiti, B.O., 2018. Business regulation and control in Nigeria. Book of readings in honour of Professor SO Otokiti, 1(2), pp.201-215.

[49] Park, P., Ergen, S.C., Fischione, C., Lu, C. and Johansson, K.H., 2017. Wireless network design for control systems: A survey. IEEE Communications Surveys & Tutorials, 20(2), pp.978-1013.

[50] Petrenko, A., 2017. Distributed Software Development Tools for Distributed Scientific Applications. Recent Progress in Parallel and Distributed Computing, p.69.

[51] Pflüger, D., Mehl, M., Valentin, J., Lindner, F., Pfander, D., Wagner, S., Graziotin, D. and Wang, Y., 2016, November. The scalability-efficiency/maintainability-portability trade-off in simulation software engineering: Examples and a preliminary systematic literature review. In 2016 Fourth International Workshop on Software Engineering for High Performance Computing in Computational Science and Engineering (SE-HPCCSE) (pp. 26-34). IEEE.

[52] Raj, P., 2018. The Hadoop ecosystem technologies and tools. In Advances in computers (Vol. 109, pp. 279-320). Elsevier.

[53] Ramakrishnan, R., Sridharan, B., Douceur, J.R., Kasturi, P., Krishnamachari-Sampath, B., Krishnamoorthy, K., Li, P., Manu, M., Michaylov, S., Ramos, R. and Sharman, N., 2017, May. Azure data lake store: a hyperscale distributed file service for big data analytics. In Proceedings of the 2017 ACM International Conference on Management of Data (pp. 51-63).

[54] Rosenbaum, S., 2017. Serverless computing in Azure with. NET. Packt Publishing Ltd.

[55] Shalev, N., 2018. Improving system security and reliability with OS help. Research Thesis.

[56] SHARMA, A., ADEKUNLE, B.I., OGEAWUCHI, J.C., ABAYOMI, A.A. and ONIFADE, O., 2019. IoT-enabled Predictive Maintenance for Mechanical Systems: Innovations in Real-time Monitoring and Operational Excellence.

[57] Shin, S., Tirukkovalluri, S.K., Tuck, J. and Solihin, Y., 2017, October. Proteus: A flexible and fast software supported hardware logging approach for nvm. In Proceedings of the 50th Annual IEEE/ACM International Symposium on Microarchitecture (pp. 178-190).

[58] Theorin, A., Bengtsson, K., Provost, J., Lieder, M., Johnsson, C., Lundholm, T. and Lennartson, B., 2017. An event-driven manufacturing information system architecture for Industry 4.0. International journal of production research, 55(5), pp.1297-1311.

[59] Thomas, G.A., Botha, R.A. and Greunen, D.V., 2018. A Virtual-Community-Centric Architecture to Support Coordination in a Large Scale Distributed Environment: A Case Study of the South African Public Sector.

[60] Vivian, J., Rao, A., Nothaft, F.A., Ketchum, C., Armstrong, J., Novak, A., Pfeil, J., Narkizian, J., Deran, A.D., Musselman-Brown, A. and Schmidt, H., 2016. Rapid and efficient analysis of 20,000 RNA-seq samples with Toil. bioRxiv, p.062497.

[61] Wang, Q., Hassan, W.U., Bates, A. and Gunter, C., 2018, February. Fear and logging in the internet of things. In Network and Distributed Systems Symposium.

How to cite this paper

Eseoghene Daniel Erigha, Ehimah Obuse, Babawale Patrick Okare, Abel Chukwuemeke Uzoka, Samuel Owoade; Noah Ayanbode "Enhancing Enterprise Software Reliability Using Retry Queues and Message Persistence in Event-Driven Cloud Environments" Iconic Research And Engineering Journals Volume 2 Issue 11 2019 Page 481-496
Eseoghene Daniel Erigha, Ehimah Obuse, Babawale Patrick Okare, Abel Chukwuemeke Uzoka, Samuel Owoade; Noah Ayanbode "Enhancing Enterprise Software Reliability Using Retry Queues and Message Persistence in Event-Driven Cloud Environments" Iconic Research And Engineering Journals, vol. 2, no. 11, May. 2019
Eseoghene Daniel Erigha, Ehimah Obuse, Babawale Patrick Okare, Abel Chukwuemeke Uzoka, Samuel Owoade; Noah Ayanbode (2019). Enhancing Enterprise Software Reliability Using Retry Queues and Message Persistence in Event-Driven Cloud Environments. Iconic Research And Engineering Journals, 2(11).
Eseoghene Daniel Erigha, Ehimah Obuse, Babawale Patrick Okare, Abel Chukwuemeke Uzoka, Samuel Owoade; Noah Ayanbode "Enhancing Enterprise Software Reliability Using Retry Queues and Message Persistence in Event-Driven Cloud Environments" Iconic Research And Engineering Journals, vol. 2, no. 11, May. 2019.
@article{1710020,
      author = {Eseoghene Daniel Erigha, Ehimah Obuse, Babawale Patrick Okare, Abel Chukwuemeke Uzoka, Samuel Owoade; Noah Ayanbode},
      title = {Enhancing Enterprise Software Reliability Using Retry Queues and Message Persistence in Event-Driven Cloud Environments},
      journal = {Iconic Research And Engineering Journals},
      year = {2019},
      volume = {2},
      number = {11},
      pages = {481-496},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1710020.pdf},
      abstract = {In an era where enterprise software systems are increasingly deployed on cloud platforms and built upon event-driven architectures, ensuring consistent reliability across distributed components becomes a critical concern. These modern architectures promote scalability and responsiveness through asynchronous communication, but they also introduce new complexities in handling transient failures, message delivery guarantees, and fault tolerance. This explores the role of retry queues and message persistence as foundational mechanisms for enhancing software reliability in such environments. Retry queues enable services to automatically attempt message processing again after initial failures, using configurable strategies such as exponential backoff, jitter, and maximum retry limits. These mechanisms help prevent message loss, reduce system downtime, and improve end-to-end transaction success rates. When integrated with dead-letter queues and observability tools, retry queues offer not only recovery but also insight into persistent system weaknesses and transient bottlenecks. Message persistence further strengthens reliability by ensuring that messages are durably stored?often across distributed logs or message brokers?until they are successfully processed or safely discarded. Leveraging technologies such as Apache Kafka, AWS SQS with Dead-Letter Queues, and Azure Service Bus, developers can implement various delivery semantics (at-least-once, exactly-once, at-most-once) suited to different application requirements. Persistence protects against system crashes, network partitions, and service restarts, thereby maintaining data integrity and continuity across the system. This synthesizes architectural best practices, cloud-native tooling, and design patterns for implementing retry logic and persistent messaging in microservice-based systems. It also highlights real-world use cases?including transactional processing, notification systems, and event sourcing?demonstrating how these reliability mechanisms can be effectively employed. Finally, the discussion explores future directions such as AI-assisted retry strategies, serverless queue orchestration, and cross-cloud persistence standards. In conclusion, retry queues and message persistence are indispensable tools for building fault-tolerant, enterprise-grade, event-driven software in dynamic cloud environments.},
      keywords = {Enterprise, Software reliability, Retry queues, Message persistence, Event-driven, Cloud environments},
      month = {May},
  }