International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1711043

1711043 Vol 5 · Issue 3 Download Paper

Markov Decision Processes with Formal Verification: Mathematical Guarantees for Safe Reinforcement Learning

Syed Khundmir Azmi

Subject area: Science,Engineering and Technology  ·  Area of research: Mathematics

Abstract

This study investigates the application of Markov Decision Processes (MDPs) in conjunction with formal verification to enhance the safety of reinforcement learning (RL) systems. The primary focus of this work is to develop approaches that provide mathematical assurances for the safe exploration of the RL environment, a significant challenge in autonomous decision-making systems. The paper examines the application of formal verification in ensuring that RL agents adhere to the specified safety restrictions while learning the optimal policies. The primary goals are to develop mathematical models that quantify safety risks and apply these models in practical settings. The observations indicate that there are substantive developments in offering verifiable safety guarantees during the exploration process, thereby reducing the chances of catastrophic failures. This work is an addition to the expanding body of safe RL by combining formal methods with MDPs, which is a new way of accomplishing reliable, safe, and efficient learning in non-trivial settings.

Keywords

Markov Decision Processes, Formal Verification, Reinforcement Learning, Safety Constraints, Mathematical Guarantees, Safe Exploration, Autonomous Decision-Making, Provable Safety, Formal Methods, Optimal Policies.

References

[1] Deshmukh, J. V., & Sriram Sankaranarayanan. (2019). Formal Techniques for Verification and Testing of Cyber-Physical Systems. Springer EBooks, 69–105. https://doi.org/10.1007/978-3-030-13050-3_4

[2] Fulton, N., & Platzer, A. (2018). Safe Reinforcement Learning via Formal Methods: Toward Safe Control Through Proof and Learning. Proceedings of the AAAI Conference on Artificial Intelligence, 32(1). https://ojs.aaai.org/index.php/AAAI/article/view/12107

[3] Grimm, T., Djones Lettnin, & Hübner, M. (2018). A Survey on Formal Verification Techniques for Safety-Critical Systems-on-Chip. Electronics, 7(6), 81–81. https://doi.org/10.3390/electronics7060081

[4] Holland, J., Kingston, L., McCarthy, C., Armstrong, E., O’Dwyer, P., Merz, F., & McConnell, M. (2021). Service Robots in the Healthcare Sector. Robotics, 10(1), 47. https://doi.org/10.3390/robotics10010047

[5] Kim, Y., Allmendinger, R., & López-Ibáñez, M. (2021). Safe Learning and Optimization Techniques: Towards a Survey of the State of the Art. Lecture Notes in Computer Science, 123–139. https://doi.org/10.1007/978-3-030-73959-1_12

[6] Li, Y., Yin, X., Wang, Z., Yao, J., Shi, X., Wu, J., Zhang, H., & Wang, Q. (2019). A Survey on Network Verification and Testing with Formal Methods: Approaches and Challenges. IEEE Communications Surveys & Tutorials, 21(1), 940–969. https://doi.org/10.1109/comst.2018.2868050

[7] Luckcuck, M., Farrell, M., Dennis, L. A., Dixon, C., & Fisher, M. (2019). Formal Specification and Verification of Autonomous Robotic Systems. ACM Computing Surveys, 52(5), 1–41. https://doi.org/10.1145/3342355

[8] Scherer, W. T., Adams, S., & Beling, P. A. (2018). On the Practical Art of State Definitions for Markov Decision Process Construction. IEEE Access, 6, 21115–21128. https://doi.org/10.1109/access.2018.2819940

[9] Sun, Z., Lin, M., Chen, W., Dai, B., Ying, P., & Zhou, Q. (2023). A case study of unavoidable accidents of autonomous vehicles. Traffic Injury Prevention, 1–6. https://doi.org/10.1080/15389588.2023.2255333

[10] Wei, Z., Xu, J., Lan, Y., Guo, J., & Cheng, X. (2017). Reinforcement Learning to Rank with Markov Decision Process. Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval. https://doi.org/10.1145/3077136.3080685

[11] Zhang, J., Cheung, B., Finn, C., Levine, S., & Jayaraman, D. (2020). Cautious Adaptation For Reinforcement Learning in Safety-Critical Settings. PMLR, 11055–11065. https://proceedings.mlr.press/v119/zhang20e.html

How to cite this paper

Syed Khundmir Azmi "Markov Decision Processes with Formal Verification: Mathematical Guarantees for Safe Reinforcement Learning" Iconic Research And Engineering Journals Volume 5 Issue 3 2021 Page 418-428
Syed Khundmir Azmi "Markov Decision Processes with Formal Verification: Mathematical Guarantees for Safe Reinforcement Learning" Iconic Research And Engineering Journals, vol. 5, no. 3, Sep. 2021
Syed Khundmir Azmi (2021). Markov Decision Processes with Formal Verification: Mathematical Guarantees for Safe Reinforcement Learning. Iconic Research And Engineering Journals, 5(3).
Syed Khundmir Azmi "Markov Decision Processes with Formal Verification: Mathematical Guarantees for Safe Reinforcement Learning" Iconic Research And Engineering Journals, vol. 5, no. 3, Sep. 2021.
@article{1711043,
      author = {Syed Khundmir Azmi},
      title = {Markov Decision Processes with Formal Verification: Mathematical Guarantees for Safe Reinforcement Learning},
      journal = {Iconic Research And Engineering Journals},
      year = {2021},
      volume = {5},
      number = {3},
      pages = {418-428},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1711043.pdf},
      abstract = {This study investigates the application of Markov Decision Processes (MDPs) in conjunction with formal verification to enhance the safety of reinforcement learning (RL) systems. The primary focus of this work is to develop approaches that provide mathematical assurances for the safe exploration of the RL environment, a significant challenge in autonomous decision-making systems. The paper examines the application of formal verification in ensuring that RL agents adhere to the specified safety restrictions while learning the optimal policies. The primary goals are to develop mathematical models that quantify safety risks and apply these models in practical settings. The observations indicate that there are substantive developments in offering verifiable safety guarantees during the exploration process, thereby reducing the chances of catastrophic failures. This work is an addition to the expanding body of safe RL by combining formal methods with MDPs, which is a new way of accomplishing reliable, safe, and efficient learning in non-trivial settings.},
      keywords = {Markov Decision Processes, Formal Verification, Reinforcement Learning, Safety Constraints, Mathematical Guarantees, Safe Exploration, Autonomous Decision-Making, Provable Safety, Formal Methods, Optimal Policies.},
      month = {September},
  }