Home / Current Issue / Paper 1722594
Multi-Agent Reinforcement Learning for Dynamic Inventory Rebalancing and Last-Mile Fulfillment Under Supply Chain Disruptions
Subject area: Science,Engineering and Technology · Area of research: Machine Learning, AI
DOI: https://doi.org/10.64388/IREV8I9-1722594
Abstract
Supply chain disruptions propagate rapidly through multi-echelon networks, and most learning-based approaches stop at prediction rather than acting on it. This paper advances from disruption forecasting toward autonomous mitigation by framing dynamic inventory rebalancing and last-mile fulfillment as a cooperative multi-agent reinforcement learning problem. Each facility in the network is an independent agent that jointly decides (i) replenishment and lateral transshipment quantities to rebalance inventory across echelons, and (ii) fulfillment assignments that reroute customer orders through available last-mile capacity when a disruption degrades primary routes. The agents are trained under a centralized-training–decentralized-execution paradigm with a graph-neural-network state encoder and a clipped proximal-policy-optimization core and are exposed during training to a stochastic disruption generator so that mitigation policies are learned proactively rather than reactively. Experiments on a three-echelon network under supplier-failure, hub-failure, and transport-disruption scenarios show that the learned policy sustains service levels, shortens recovery time, and reduces total disruption cost relative to base-stock and single-agent baselines. The results demonstrate that prediction must be coupled with autonomous optimization to deliver practical resilience.
Keywords
multi-agent reinforcement learning, inventory rebalancing, last-mile fulfillment, supply chain disruptions, resilience optimization, autonomous decision-making.
References
[1] S. D. A. PAULO, “Beyond Forecasting: Deep Reinforcement Learning for Proactive Supply Chain Resilience,” Zenodo (CERN European Organization for Nuclear Research) , Nov. 2025, doi: 10.5281/zenodo.17644828.
[2] B. T. K. Uyen and B. T. Hieu, “Toward Autonomous Supply Chains: A Deep Reinforcement Learning Framework,” International Journal of Advanced Multidisciplinary Research and Studies , vol. 6, no. 2, pp. 301–312, Mar. 2026, doi: 10.62225/2583049x.2026.6.2.5959.
[3] M. A. M. A. Mousa, D. van de Berg, N. Kotecha, E. A. del Rio‐Chanona, and M. Mowbray, “An analysis of multi-agent reinforcement learning for decentralized inventory control systems,” Computers & Chemical Engineering , vol. 188, pp. 108783–108783, Jun. 2024, doi: 10.1016/j.compchemeng.2024.108783.
[4] M. Khirwar, K. S. Gurumoorthy, A. A. Jain, and S. Manchenahally, “Cooperative Multi-Agent Reinforcement Learning for Inventory Management,” Apr. 18, 2023, Cornell University . doi: 10.48550/arxiv.2304.08769.
[5] X. Yang et al. , “A Versatile Multi-Agent Reinforcement Learning Benchmark for Inventory Management,” Jun. 13, 2023, Cornell University . doi: 10.48550/arxiv.2306.07542.
[6] F. Stranieri, F. Stella, and C. Kouki, “Performance of deep reinforcement learning algorithms in two-echelon inventory control systems,” International Journal of Production Research , vol. 62, no. 17, pp. 6211–6226, Mar. 2024, doi: 10.1080/00207543.2024.2311180.
[7] K. Geevers, L. van Hezewijk, and M. Mes, “Multi-echelon inventory optimization using deep reinforcement learning,” Central European Journal of Operations Research , vol. 32, no. 3, pp. 653–683, Jul. 2023, doi: 10.1007/s10100-023-00872-2.
[8] N. Kotecha and E. A. del Rio‐Chanona, “Leveraging graph neural networks and multi-agent reinforcement learning for inventory control in supply chains,” Computers & Chemical Engineering , vol. 199, pp. 109111–109111, Apr. 2025, doi: 10.1016/j.compchemeng.2025.109111.
[9] Y. Zhao and C. Hayes, “Hierarchical Multi-Agent Reinforcement Learning for Dynamic Inventory Allocation with Demand Uncertainty,” Multidisciplinary Research in Computing Information Systems , vol. 5, no. 10, pp. 850–872, Dec. 2025, doi: 10.71465/mrcis153.
[10] G. Wu, M. Á. de C. Servia, and M. Mowbray, “Distributional reinforcement learning for inventory management in multi-echelon supply chains,” Digital Chemical Engineering , vol. 6, pp. 100073–100073, Dec. 2022, doi: 10.1016/j.dche.2022.100073.
[11] M. Silva, J. P. Pedroso, and A. Viana, “Deep reinforcement learning for stochastic last-mile delivery with crowdshipping,” Econstor (Econstor) , vol. 12, pp. 100105–100105, Jan. 2023, doi: 10.1016/j.ejtl.2023.100105.
[12] M. Silva and J. P. Pedroso, “Deep Reinforcement Learning for Crowdshipping Last-Mile Delivery with Endogenous Uncertainty,” Mathematics , vol. 10, no. 20, pp. 3902–3902, Oct. 2022, doi: 10.3390/math10203902.
[13] R. Auad, A. L. Erera, and M. Savelsbergh, “Dynamic Courier Capacity Acquisition in Rapid Delivery Systems: A Deep Q-Learning Approach,” Transportation Science , vol. 58, no. 1, pp. 67–93, Dec. 2023, doi: 10.1287/trsc.2022.0042.
[14] Y. Liu, L. Tang, Z. He, Y. Zhao, J. Ma, and Z. Duan, “Vehicle Rebalancing Under Supply-Demand Uncertainty: A Robust Multi-Agent Reinforcement Learning Framework,” in Advances in transdisciplinary engineering , IOS Press, 2025. doi: 10.3233/atde250450.
[15] J. Xi, F. Zhu, P. Ye, Y. Lv, G. Xiong, and F. Wang, “Auxiliary Network Enhanced Hierarchical Graph Reinforcement Learning for Vehicle Repositioning,” IEEE Transactions on Intelligent Transportation Systems , vol. 25, no. 9, pp. 11563–11575, Apr. 2024, doi: 10.1109/tits.2024.3383720.
[16] Z. Yu and M. Hu, “Deep Reinforcement Learning With Graph Representation for Vehicle Repositioning,” IEEE Transactions on Intelligent Transportation Systems , vol. 23, no. 8, pp. 13094–13107, Oct. 2021, doi: 10.1109/tits.2021.3119662.
[17] W. J. Tan, W. Cai, and A. N. Zhang, “Structural-aware simulation analysis of supply chain resilience,” International Journal of Production Research , vol. 58, no. 17, pp. 5175–5195, Dec. 2019, doi: 10.1080/00207543.2019.1705421.
[18] D. Ivanov, B. Sokolov, and A. Dolgui, “The Ripple effect in supply chains: trade-off ‘efficiency-flexibility-resilience’ in disruption management,” HAL (Le Centre pour la Communication Scientifique Directe) , vol. 52, no. 7, pp. 2154–2172, Nov. 2013, doi: 10.1080/00207543.2013.858836.
[19] D. Ivanov and A. Dolgui, “Stress testing supply chains and creating viable ecosystems,” Operations Management Research , vol. 15, pp. 475–486, May 2021, doi: 10.1007/s12063-021-00194-z.
[20] A. Dolgui, D. Ivanov, and B. Sokolov, “Ripple effect in the supply chain: an analysis and recent literature,” International Journal of Production Research , vol. 56, pp. 414–430, Oct. 2017, doi: 10.1080/00207543.2017.1387680.
[21] M. Fattahi, K. Govindan, and R. Maihami, “Stochastic optimization of disruption-driven supply chain network design with a new resilience metric,” University of Southern Denmark Research Portal (University of Southern Denmark) , vol. 230, pp. 107755–107755, Apr. 2020, doi: 10.1016/j.ijpe.2020.107755.
[22] K. A. Loganathan and A. Chinnaraju, “AI-Native Supply Chain Resilience: A Multimodal Architecture for Predictive Intelligence, Optimization, and Real-Time Decision-Making,” GSC Advanced Research and Reviews , vol. 25, no. 3, pp. 78–126, Dec. 2025, doi: 10.30574/gscarr.2025.25.3.0372.
[23] M. Luo et al. , “Fleet Rebalancing for Expanding Shared e-Mobility Systems: A Multi-Agent Deep Reinforcement Learning Approach,” Warwick Research Archive Portal (University of Warwick) , vol. 24, no. 4, pp. 3868–3881, Jan. 2023, doi: 10.1109/tits.2022.3233422.
[24] F. Stranieri, E. Fadda, and F. Stella, “Combining deep reinforcement learning and multi-stage stochastic programming to address the supply chain inventory management problem,” BOA (University of Milano-Bicocca) , vol. 268, pp. 109099–109099, Nov. 2023, doi: 10.1016/j.ijpe.2023.109099.
How to cite this paper
@article{1722594,
author = {Sohail Sayed, Nauman Sayed},
title = {Multi-Agent Reinforcement Learning for Dynamic Inventory Rebalancing and Last-Mile Fulfillment Under Supply Chain Disruptions},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {8},
number = {9},
pages = {2070-2078},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722594.pdf},
abstract = {Supply chain disruptions propagate rapidly through multi-echelon networks, and most learning-based approaches stop at prediction rather than acting on it. This paper advances from disruption forecasting toward autonomous mitigation by framing dynamic inventory rebalancing and last-mile fulfillment as a cooperative multi-agent reinforcement learning problem. Each facility in the network is an independent agent that jointly decides (i) replenishment and lateral transshipment quantities to rebalance inventory across echelons, and (ii) fulfillment assignments that reroute customer orders through available last-mile capacity when a disruption degrades primary routes. The agents are trained under a centralized-training–decentralized-execution paradigm with a graph-neural-network state encoder and a clipped proximal-policy-optimization core and are exposed during training to a stochastic disruption generator so that mitigation policies are learned proactively rather than reactively. Experiments on a three-echelon network under supplier-failure, hub-failure, and transport-disruption scenarios show that the learned policy sustains service levels, shortens recovery time, and reduces total disruption cost relative to base-stock and single-agent baselines. The results demonstrate that prediction must be coupled with autonomous optimization to deliver practical resilience.},
keywords = {multi-agent reinforcement learning, inventory rebalancing, last-mile fulfillment, supply chain disruptions, resilience optimization, autonomous decision-making.},
month = {March},
doi = {https://doi.org/10.64388/IREV8I9-1722594}
}