International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1722951

1722951 Vol 9 · Issue 3 Download Paper

Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support

Sohail Sayed Nauman Sayed

Subject area: Science,Engineering and Technology  ·  Area of research: Machine Learning, AI

DOI: https://doi.org/10.64388/IREV9I3-1722951

Abstract

Large language models are prone to hallucination, generating plausible yet nonfactual content — a phenomenon that raises significant concerns over their reliability in real-world systems (Huang et al., 2024). In supply chains, such hallucinations can lead to erroneous demand forecasts or misinterpretation of supply chain relationships, potentially resulting in operational disruptions and financial losses: a generative model might incorrectly predict a demand surge based on fabricated trends, leading to overproduction and increased inventory costs (Ge & Brintrup, 2024). Yet the deployment pressure is real: LLMs are already used for supplier risk assessment, procurement contracting, inventory optimization, and disruption response, where one unsupported number can cascade through a network. This paper argues that trustworthiness in high-stakes supply chain decision support is an engineered property, not an emergent one: it requires hallucination awareness (measurement and detection), grounding (retrieval and knowledge-graph anchoring), verification (self-consistency, semantic entropy, conformal guarantees), and governance (guardrails, abstention, and human-final authority). We propose HALO-SC, a hallucination-aware framework coupling (i) request triage with risk-tier classification, (ii) retrieval and knowledge-graph grounding, (iii) generation with uncertainty quantification via self-consistency and semantic entropy, (iv) multi-signal verification including LLM-as-a-judge with known-bias compensation, (v) conformal abstention with bounded hallucination rate, (vi) deterministic guardrail enforcement with GO/HOLD/NO-GO decision states, and (vii) audit-ready decision lineage. The framework consolidates the reported evidence envelope: retrieval-augmented pipelines reduce hallucination rates to 1.5% on average versus 8.3% for non-RAG baselines — an 81.9% relative reduction — with retrieval precision above 91% across product domains (Kasarapu, 2026); hallucination rate falls from 6.8% at 10,000 documents to 1.2% at 100,000 documents, establishing knowledge-base enrichment as the dominant reliability lever (Kasarapu, 2026); conformal abstention procedures bound the hallucination rate at a user-specified error level with rigorous theoretical guarantees (Yadkori et al., 2024); semantic entropy detects confabulations in free-form generation (Farquhar et al., 2024); governed LLM-based optimization reduces unsafe decision outputs by 45% while sustaining drift-detection accuracy above 92% (Jingar, 2023); decision authority frameworks formalize GO/HOLD/NO-GO governance states with human-final authority enforcement (KALAFATOGLU, 2025, 2026); and human–AI collaboration in which humans specify the optimization algorithm achieves statistically optimal and stable outcomes under supply chain disruption (Wu, 2026). The paper argues that the trustworthiness ladder — grounded, measured, verified, abstaining, governed — is what separates LLM assistants from LLM advisors in commerce.

Keywords

Hallucination, large language models, trustworthy AI, uncertainty quantification, conformal prediction, semantic entropy, supply chain decision support, guardrails, abstention, human-in-the-loop, knowledge graphs.

How to cite this paper

Sohail Sayed, Nauman Sayed "Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support" Iconic Research And Engineering Journals Volume 9 Issue 3 2025 Page 2341-2357 https://doi.org/10.64388/IREV9I3-1722951
Sohail Sayed, Nauman Sayed "Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support" Iconic Research And Engineering Journals, vol. 9, no. 3, Sep. 2025, doi: https://doi.org/10.64388/IREV9I3-1722951
Sohail Sayed, Nauman Sayed (2025). Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support. Iconic Research And Engineering Journals, 9(3). doi: https://doi.org/10.64388/IREV9I3-1722951
Sohail Sayed, Nauman Sayed "Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support" Iconic Research And Engineering Journals, vol. 9, no. 3, Sep. 2025. Crossref, https://doi.org/10.64388/IREV9I3-1722951
@article{1722951,
      author = {Sohail Sayed, Nauman Sayed},
      title = {Hallucination-Aware and Trustworthy LLMs for High-Stakes Supply Chain Decision Support},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {3},
      pages = {2341-2357},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722951.pdf},
      abstract = {Large language models are prone to hallucination, generating plausible yet nonfactual content — a phenomenon that raises significant concerns over their reliability in real-world systems (Huang et al., 2024). In supply chains, such hallucinations can lead to erroneous demand forecasts or misinterpretation of supply chain relationships, potentially resulting in operational disruptions and financial losses: a generative model might incorrectly predict a demand surge based on fabricated trends, leading to overproduction and increased inventory costs (Ge & Brintrup, 2024). Yet the deployment pressure is real: LLMs are already used for supplier risk assessment, procurement contracting, inventory optimization, and disruption response, where one unsupported number can cascade through a network. This paper argues that trustworthiness in high-stakes supply chain decision support is an engineered property, not an emergent one: it requires hallucination awareness (measurement and detection), grounding (retrieval and knowledge-graph anchoring), verification (self-consistency, semantic entropy, conformal guarantees), and governance (guardrails, abstention, and human-final authority). We propose HALO-SC, a hallucination-aware framework coupling (i) request triage with risk-tier classification, (ii) retrieval and knowledge-graph grounding, (iii) generation with uncertainty quantification via self-consistency and semantic entropy, (iv) multi-signal verification including LLM-as-a-judge with known-bias compensation, (v) conformal abstention with bounded hallucination rate, (vi) deterministic guardrail enforcement with GO/HOLD/NO-GO decision states, and (vii) audit-ready decision lineage. The framework consolidates the reported evidence envelope: retrieval-augmented pipelines reduce hallucination rates to 1.5% on average versus 8.3% for non-RAG baselines — an 81.9% relative reduction — with retrieval precision above 91% across product domains (Kasarapu, 2026); hallucination rate falls from 6.8% at 10,000 documents to 1.2% at 100,000 documents, establishing knowledge-base enrichment as the dominant reliability lever (Kasarapu, 2026); conformal abstention procedures bound the hallucination rate at a user-specified error level with rigorous theoretical guarantees (Yadkori et al., 2024); semantic entropy detects confabulations in free-form generation (Farquhar et al., 2024); governed LLM-based optimization reduces unsafe decision outputs by 45% while sustaining drift-detection accuracy above 92% (Jingar, 2023); decision authority frameworks formalize GO/HOLD/NO-GO governance states with human-final authority enforcement (KALAFATOGLU, 2025, 2026); and human–AI collaboration in which humans specify the optimization algorithm achieves statistically optimal and stable outcomes under supply chain disruption (Wu, 2026). The paper argues that the trustworthiness ladder — grounded, measured, verified, abstaining, governed — is what separates LLM assistants from LLM advisors in commerce.},
      keywords = {Hallucination, large language models, trustworthy AI, uncertainty quantification, conformal prediction, semantic entropy, supply chain decision support, guardrails, abstention, human-in-the-loop, knowledge graphs.},
      month = {September},
      doi = {https://doi.org/10.64388/IREV9I3-1722951}
  }