International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1722951

1722951 Vol 9 · Issue 3 Download Paper

Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support

Sohail Sayed Nauman Sayed

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning, AI

DOI: 10.64388/IREV9I3-1722951

Abstract

Large language models are prone to hallucination, generating plausible yet nonfactual content — a phenomenon that raises significant concerns over their reliability in real-world systems (Huang et al., 2024). In supply chains, such hallucinations can lead to erroneous demand forecasts or misinterpretation of supply chain relationships, potentially resulting in operational disruptions and financial losses: a generative model might incorrectly predict a demand surge based on fabricated trends, leading to overproduction and increased inventory costs (Ge & Brintrup, 2024). Yet the deployment pressure is real: LLMs are already used for supplier risk assessment, procurement contracting, inventory optimization, and disruption response, where one unsupported number can cascade through a network. This paper argues that trustworthiness in high-stakes supply chain decision support is an engineered property, not an emergent one: it requires hallucination awareness (measurement and detection), grounding (retrieval and knowledge-graph anchoring), verification (self-consistency, semantic entropy, conformal guarantees), and governance (guardrails, abstention, and human-final authority). We propose HALO-SC, a hallucination-aware framework coupling (i) request triage with risk-tier classification, (ii) retrieval and knowledge-graph grounding, (iii) generation with uncertainty quantification via self-consistency and semantic entropy, (iv) multi-signal verification including LLM-as-a-judge with known-bias compensation, (v) conformal abstention with bounded hallucination rate, (vi) deterministic guardrail enforcement with GO/HOLD/NO-GO decision states, and (vii) audit-ready decision lineage. The framework consolidates the reported evidence envelope: retrieval-augmented pipelines reduce hallucination rates to 1.5% on average versus 8.3% for non-RAG baselines — an 81.9% relative reduction — with retrieval precision above 91% across product domains (Kasarapu, 2026); hallucination rate falls from 6.8% at 10,000 documents to 1.2% at 100,000 documents, establishing knowledge-base enrichment as the dominant reliability lever (Kasarapu, 2026); conformal abstention procedures bound the hallucination rate at a user-specified error level with rigorous theoretical guarantees (Yadkori et al., 2024); semantic entropy detects confabulations in free-form generation (Farquhar et al., 2024); governed LLM-based optimization reduces unsafe decision outputs by 45% while sustaining drift-detection accuracy above 92% (Jingar, 2023); decision authority frameworks formalize GO/HOLD/NO-GO governance states with human-final authority enforcement (KALAFATOGLU, 2025, 2026); and human–AI collaboration in which humans specify the optimization algorithm achieves statistically optimal and stable outcomes under supply chain disruption (Wu, 2026). The paper argues that the trustworthiness ladder — grounded, measured, verified, abstaining, governed — is what separates LLM assistants from LLM advisors in commerce.

Keywords

Hallucination, large language models, trustworthy AI, uncertainty quantification, conformal prediction, semantic entropy, supply chain decision support, guardrails, abstention, human-in-the-loop, knowledge graphs.

References

[1] Akubilla, J., Somoye, O. I., Abiodun, F., & Serifat, O. A. (2025). The Role of Explainable AI in Enhancing Trust and Transparency in Supply Chain Risk Mitigation. International Journal of Multidisciplinary Research and Growth Evaluation, 6(3), 367–377. [Crossref]

[2] AlMahri, S., Xu, L., & Brintrup, A. (2026). Automating Supply Chain Disruption Monitoring via an Agentic AI Approach. ArXiv.Org. [Crossref]

[3] Baryannis, G., Dani, S., & Antoniou, G. (2019). Predicting supply chain risks using machine learning: The trade-off between performance and interpretability. Future Generation Computer Systems, 101, 993–1004. [Crossref]

[4] Brandtner, P., & Hofer, F. (2026). Enhancing Procurement Processes in Supply Chain Management with Large Language Models. Procedia Computer Science, 278, 308–315. [Crossref]

[5] Campos, M., Farinhas, A., Zerva, C., Figueiredo, M. A. T., & Martins, A. F. T. (2024). Conformal Prediction for Natural Language Processing: A Survey. Transactions of the Association for Computational Linguistics, 12, 1497–1516. [Crossref]

[6] Chakraborty, N., Ornik, M., & Driggs-Campbell, K. (2025). Hallucination Detection in Foundation Models for Decision-Making: A Flexible Definition and Review of the State of the Art. In arXiv (Cornell University) (Vol. 57, Issue 7, pp. 1–35). Cornell University. [Crossref]

[7] Chen, X., Wanigarathna, Y. R., Chattopadhyay, A., Tarkoma, S., & Rao, A. (2025). A Pedagogical Approach for Evaluating AIOps. 40–49. [Crossref]

[8] Cherian, J. J., Gibbs, I., & Candès, E. J. (2024). Large language model validity via enhanced conformal prediction methods. In arXiv (Cornell University). Cornell University. [Crossref]

[9] Cornacchia, A., Alabdulaal, I., Saghier, I., Mirdad, A., Fayoumi, O., & Canini, M. (2025). Between Promise and Pain: The Reality of Automating Failure Analysis in Microservices with LLMs. 155–167. [Crossref]

[10] Cui, Y., Zeng, Y., Yan, J., Lin, K. L., Ji, K. Y., Zeng, J., Zhang, S., Luo, X., Su, B., Shen, C., & Yu, J. (2026). CSCBench: A PVC Diagnostic Benchmark for Commodity Supply Chain Reasoning. In arXiv (Cornell University). Cornell University. [Crossref]

[11] Davenport. (2026). GUARDRAIL-CENTRIC FINE-TUNING FOR DETERMINISTIC DECISION SYSTEMS. Zenodo (CERN European Organization for Nuclear Research). [Crossref]

[12] Dehan, M. F. Z., Anzum, K. Md. T., Disha, J. F., Masum, MD. M. R., & Mahmud, I. (2025). A Systematic Review on the Applications of How Explainable Artificial Intelligence can be Applied in the Supply Chain Management. [Crossref]

[13] Farquhar, S., Kossen, J., Kuhn, L., & Gal, Y. (2024). Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017), 625–630. [Crossref]

[14] Ge, Z., & Brintrup, A. (2024). Enhancing Supply Chain Visibility with Generative AI: An Exploratory Case Study on Relationship Prediction in Knowledge Graphs. In arXiv (Cornell University). Cornell University. [Crossref]

[15] Guan, S., Liu, Y., & Cao, L. (2026). SupChain-Bench: Benchmarking Large Language Models for Real-World Supply Chain Management. arXiv (Cornell University). [Crossref]

[16] Gumaan, E. (2025). Theoretical Foundations and Mitigation of Hallucination in Large Language Models. In ArXiv.org. [Crossref]

[17] Huang, L., Yu, W., Ma, W., Zhong, W., Feng, Z., Wang, H., Chen, Q., Peng, W., Feng, X., Qin, B., & Liu, T. (2024). A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ACM Transactions on Information Systems, 43(2), 1–55. [Crossref]

[18] Iruku, V. M. (2025). Multi-dimensional XAI Framework Revealing Critical Supply Chain Vulnerability Drivers. World Journal of Advanced Engineering Technology and Sciences, 15(3), 2141–2152. [Crossref]

[19] Jannelli, V., Schöpf, S., Bickel, M., Netland, T. H., & Brintrup, A. (2025). Agentic LLMs in the supply chain: towards autonomous multi-agent consensus-seeking. International Journal of Production Research, 1–31. [Crossref]

[20] Jingar, N. K. (2023). Ensuring Safety, Accountability, and Drift Resistance in LLM-Based Supply Chain Optimization. International Journal of Scientific Research in Science Engineering and Technology, 472. [Crossref]

[21] KALAFATOGLU, Y. (2025). Human-Final Decision Authority in Artificial Intelligence: A Deterministic and Auditable Governance Architecture. Zenodo (CERN European Organization for Nuclear Research). [Crossref]

[22] KALAFATOGLU, Y. (2026). A Deterministic Decision Authority Framework for Governance-Grade AI Systems. In Zenodo (CERN European Organization for Nuclear Research). European Organization for Nuclear Research. [Crossref]

[23] Kang, S., Bakman, Y. F., Yaldiz, D. N., Buyukates, B., & Avestimehr, S. (2025). Uncertainty Quantification for Hallucination Detection in Large Language Models: Foundations, Methodology, and Future Directions. In ArXiv.org. [Crossref]

[24] Kasarapu, B. C. (2026). Generative AI-Enabled Micro-Frontend Framework for Scalable and Intelligent Enterprise Retail Applications. International Journal of Computer Information Systems and Industrial Management Applications, 18, 297–309. [Crossref]

[25] Kaur, D., Uslu, S., Rittichier, K. J., & Durresi, A. (2022). Trustworthy Artificial Intelligence: A Review. ACM Computing Surveys, 55(2), 1–38. [Crossref]

[26] Li, X., Wang, S., Zeng, S., Wu, Y., & Yang, Y. (2024). A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges. Vicinagearth., 1(1). [Crossref]

[27] Liu, Y., Pei, C., Xu, L., Chen, B., Sun, M., Zhang, Z. L., Sun, Y., Zhang, S., Wang, K., Zhang, H., Li, J., Xie, G., wen, xiaohui, Nie, X., Ma, M., & Pei, D. (2025). OpsEval: A Comprehensive Benchmark Suite for Evaluating Large Language Models’ Capability in IT Operations Domain. 503–513. [Crossref]

[28] Liu, Z. (2025). Judicial Hallucination: Investigating and Understanding the Unreliability of LLM-as-a-Judge. Applied and Computational Engineering, 203(1), 105–113. [Crossref]

[29] Padhy, A. K., Patel, T., Soni, V., Shivam, S., Thokala, G. B., & Vulugundam, B. (2026). Machine Learning-Based Fault Prediction in Large- Scale Distributed Systems. 1–6. [Crossref]

[30] R., K. S., Ojha, D., Kaur, P., Mahto, R. V., & Dhir, A. (2024). Explainable artificial intelligence and agile decision-making in supply chain cyber resilience. University of North Texas Digital Library (University of North Texas), 180, 114194. [Crossref]

[31] Shorinwa, O., Mei, Z., Lidard, J., Ren, A. Z., & Majumdar, A. (2025). A Survey on Uncertainty Quantification of Large Language Models: Taxonomy, Open Research Challenges, and Future Directions. ACM Computing Surveys, 58(3), 1–38. [Crossref]

[32] Syed, T. A., Belgaum, M. R., Jan, S., Khan, A. A., & Alqahtani, S. S. (2025). Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation. ArXiv.Org. [Crossref]

[33] USA, S. S. E., Dallas, Texas, & Natta, P. K. (2024). Designing Trustworthy AI Systems for Mission-Critical Enterprise Operations. International Journal of Future Innovative Science and Technology, 7(6). [Crossref]

[34] Wang, Y., Pan, Y., Su, Z., Deng, Y., Zhao, Q., Du, L., Luan, T. H., Kang, J., & Niyato, D. (2025). Large Model-Based Agents: State-of-the-Art, Cooperation Paradigms, Security and Privacy, and Future Trends. IEEE Communications Surveys & Tutorials, 28, 1906–1949. [Crossref]

[35] Wasserkrug, S., Boussioux, L., Hertog, D. den, Mirzazadeh, F., Birbil, Ş. İ., Kurtz, J., & Maragno, D. (2025). Enhancing Decision Making Through the Integration of Large Language Models and Operations Research Optimization. Proceedings of the AAAI Conference on Artificial Intelligence, 39(27), 28643–28650. [Crossref]

[36] Wu, R. (2026). Inventory optimization under supply chain disruptions: Leveraging large language models for human-AI collaborative decision-making. Journal of King Saud University - Computer and Information Sciences, 38(2). [Crossref]

[37] Yadkori, Y. A., Kuzborskij, I., Stutz, D., György, A., Fisch, A., Douc, R., Beloshapka, I., Weng, W.-H., Yang, Y.-Y., Szepesvári, C., Cemgil, A. T., & Tomasev, N. (2024). Mitigating LLM Hallucinations via Conformal Abstention. In arXiv (Cornell University). Cornell University. [Crossref]

[38] Zhang, G. (2025). Large Language Model enabled Mathematical Modeling. In ArXiv.org. [Crossref]

[39] Zhang, Y., Li, Y., Cui, L., Deng, C., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., Chen, Y., Wang, L., Luu, A. T., Bi, W., Shi, F., & Shi, S. (2025). 🧜Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. Computational Linguistics, 51(4), 1373–1418. [Crossref]

How to cite this paper

Sohail Sayed, Nauman Sayed "Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support" Iconic Research And Engineering Journals Volume 9 Issue 3 2025 Page 2341-2357 https://doi.org/10.64388/IREV9I3-1722951
Sohail Sayed, Nauman Sayed "Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support" Iconic Research And Engineering Journals, vol. 9, no. 3, Sep. 2025, doi: https://doi.org/10.64388/IREV9I3-1722951
Sohail Sayed, Nauman Sayed (2025). Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support. Iconic Research And Engineering Journals, 9(3). doi: https://doi.org/10.64388/IREV9I3-1722951
Sohail Sayed, Nauman Sayed "Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support" Iconic Research And Engineering Journals, vol. 9, no. 3, Sep. 2025. Crossref, https://doi.org/10.64388/IREV9I3-1722951
@article{1722951,
      author = {Sohail Sayed, Nauman Sayed},
      title = {Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {3},
      pages = {2341-2357},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722951.pdf},
      abstract = {Large language models are prone to hallucination, generating plausible yet nonfactual content — a phenomenon that raises significant concerns over their reliability in real-world systems (Huang et al., 2024). In supply chains, such hallucinations can lead to erroneous demand forecasts or misinterpretation of supply chain relationships, potentially resulting in operational disruptions and financial losses: a generative model might incorrectly predict a demand surge based on fabricated trends, leading to overproduction and increased inventory costs (Ge & Brintrup, 2024). Yet the deployment pressure is real: LLMs are already used for supplier risk assessment, procurement contracting, inventory optimization, and disruption response, where one unsupported number can cascade through a network. This paper argues that trustworthiness in high-stakes supply chain decision support is an engineered property, not an emergent one: it requires hallucination awareness (measurement and detection), grounding (retrieval and knowledge-graph anchoring), verification (self-consistency, semantic entropy, conformal guarantees), and governance (guardrails, abstention, and human-final authority). We propose HALO-SC, a hallucination-aware framework coupling (i) request triage with risk-tier classification, (ii) retrieval and knowledge-graph grounding, (iii) generation with uncertainty quantification via self-consistency and semantic entropy, (iv) multi-signal verification including LLM-as-a-judge with known-bias compensation, (v) conformal abstention with bounded hallucination rate, (vi) deterministic guardrail enforcement with GO/HOLD/NO-GO decision states, and (vii) audit-ready decision lineage. The framework consolidates the reported evidence envelope: retrieval-augmented pipelines reduce hallucination rates to 1.5% on average versus 8.3% for non-RAG baselines — an 81.9% relative reduction — with retrieval precision above 91% across product domains (Kasarapu, 2026); hallucination rate falls from 6.8% at 10,000 documents to 1.2% at 100,000 documents, establishing knowledge-base enrichment as the dominant reliability lever (Kasarapu, 2026); conformal abstention procedures bound the hallucination rate at a user-specified error level with rigorous theoretical guarantees (Yadkori et al., 2024); semantic entropy detects confabulations in free-form generation (Farquhar et al., 2024); governed LLM-based optimization reduces unsafe decision outputs by 45% while sustaining drift-detection accuracy above 92% (Jingar, 2023); decision authority frameworks formalize GO/HOLD/NO-GO governance states with human-final authority enforcement (KALAFATOGLU, 2025, 2026); and human–AI collaboration in which humans specify the optimization algorithm achieves statistically optimal and stable outcomes under supply chain disruption (Wu, 2026). The paper argues that the trustworthiness ladder — grounded, measured, verified, abstaining, governed — is what separates LLM assistants from LLM advisors in commerce.},
      keywords = {Hallucination, large language models, trustworthy AI, uncertainty quantification, conformal prediction, semantic entropy, supply chain decision support, guardrails, abstention, human-in-the-loop, knowledge graphs.},
      month = {September},
      doi = {https://doi.org/10.64388/IREV9I3-1722951}
  }