International Peer-Reviewed JournalOpen AccessISSN 2456-8880
irejournals@gmail.com+91-7433024337

Home / Current Issue / Paper 1717980

1717980 Vol 9 · Issue 11 Download Paper

A Multi-Metric Evaluation Perspective on Hallucination Detection in Low-Resource Governance Documents

Pranjal Gahlot Prof. Rakshitha BS

Subject area: Science,Engineering and Technology  ·  Area of research: AI, NLP, LLMs

DOI: https://doi.org/10.64388/IREV9I11-1717980

Abstract

The rapid advancement of Large Language Models (LLMs) has significantly improved natural language processing applications across domains such as governance, healthcare, legal analysis, and public information systems. Despite these advancements, LLMs frequently generate hallucinated outputs, where responses appear plausible but contain incorrect or fabricated information. This issue poses serious risks in governance-related applications, where inaccurate information can influence policy interpretation, administrative decision-making, and public trust. Existing studies have proposed several approaches to address hallucinations, including semantic entropy–based detection, benchmark evaluation frameworks, and adversarial testing methods. However, the literature indicates that current solutions remain fragmented and often focus on isolated aspects such as model performance, dataset construction, or benchmark capability rather than comprehensive reliability assessment. This literature review examines recent research on hallucination detection, multilingual and low-resource natural language processing, and evaluation frameworks for LLM reliability. The reviewed studies highlight key challenges, including the lack of multilingual hallucination evaluation, insufficient harm-oriented risk assessment, and limited adversarial robustness testing in governance contexts. Furthermore, existing benchmarks often measure task accuracy rather than factual reliability or societal impact. Based on the analysis of the literature, this review identifies major methodological and contextual gaps and proposes the need for an integrated evaluation framework combining meaning level hallucination detection, harm aware risk modeling, and multilingual robustness assessment. Such an approach could improve the reliability and safety of LLM systems deployed in governance and public service environments.

Keywords

Large Language Models, Hallucination Detection, Semantic Entropy, Multilingual NLP, Low-Resource Languages, Governance AI, Adversarial Prompting, Benchmark Evaluation, AI Reliability, Natural Language Processing.

References

[1] S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal, “Detecting hallucinations in large language models using semantic entropy,” Nature, vol. 630, pp. 625–630, 2024, doi: 10.1038/s41586-024-07421-0.

[2] Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, Art. no. 248, 2023, doi: 10.1145/3571730.

[3] L. Huang et al., “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” ACM Transactions on Information Systems, vol. 43, no. 2, Art. no. 42, 2025, doi: 10.1145/3703155.

[4] S. Liu et al., “A hallucination detection and mitigation framework for faithful text summarization using large language models,” Scientific Reports, vol. 16, Art. no. 1374, 2026, doi: 10.1038/s41598-025-31075-1.

[5] E. Asgari et al., “A framework to assess clinical safety and hallucination rates of large language models for medical text summarisation,” npj Digital Medicine, vol. 8, Art. no. 274, 2025, doi: 10.1038/s41746-025-01670-7.

[6] M. Omar et al., “Large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support,” Communications Medicine, vol. 5, Art. no. 330, 2025, doi: 10.1038/s43856-025-01021-3.

[7] R. Massenon et al., “‘My AI is lying to me’: User-reported LLM hallucinations in AI mobile apps reviews,” Scientific Reports, vol. 15, Art. no. 30397, 2025, doi: 10.1038/s41598-025-15416-8.

[8] A. Alansari and H. Luqman, “LLM hallucination: A comprehensive survey,” arXiv preprint, arXiv:2510.06265, 2025.

[9] P. Roy, “Deep ensemble network for sentiment analysis in bi-lingual low-resource languages,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 1, Art. no. 8, 2024, doi: 10.1145/3600229.

[10] P. Pakray et al., “Natural language processing applications for low-resource languages,” Natural Language Processing, vol. 31, pp. 183–197, 2025, doi: 10.1017/nlp.2024.33.

[11] I. Chalkidis et al., “LexGLUE: A benchmark dataset for legal language understanding in English,” in Proc. ACL, 2022, pp. 4310–4330.

[12] F. Ariai et al., “Natural language processing for the legal domain: A survey of tasks, datasets, models, and challenges,” ACM Computing Surveys, vol. 58, no. 6, Art. no. 163, 2025, doi: 10.1145/3777009.

[13] D. Yadav et al., “Cross-lingual named entity recognition for low-resource languages,” in Proc. ACL Workshop on Multilingual Representation Learning, 2024, pp. 167–174.

[14] S. Maddu and V. Sanapala, “A survey on NLP tasks, resources, and techniques for low-resource Telugu-English code-mixed text,” ACM Transactions on Asian and Low-Resource Language Information Processing, 2024, doi: 10.1145/3695766.

[15] M. Omar et al., “Large language models are highly vulnerable to adversarial hallucination,” medRxiv preprint, doi: 10.1101/2025.03.18.25324184, 2025.

[16] J. Zhang et al., “Neural machine translation for low-resource languages: A survey,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 6, Art. no. 80, 2024, doi: 10.1145/3665244.

[17] D. Sulistyo et al., “Pivoted low-resource multilingual translation with named entity optimization,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 24, no. 5, 2025, doi: 10.1145/3727876.

[18] B. Wanjawa et al., “KenSwQuAD—A question answering dataset for Swahili low-resource language,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 22, no. 4, Art. no. 113, 2023, doi: 10.1145/3578553.

[19] M. Munaf et al., “Low-resource summarization using pre-trained language models,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 23, no. 10, Art. no. 141, 2024, doi: 10.1145/3675780.

[20] E. Ramdinmawii and S. Nath, “Resource building and classification of Mizo folk songs,” Natural Language Processing, vol. 31, pp. 655–673, 2024, doi: 10.1017/nlp.2024.23.

[21] A. Üstün et al., “Aya model: An instruction finetuned open-access multilingual language model,” in Proc. ACL, 2024, pp. 15894–15939.

[22] BigScience Workshop, “BLOOM: A 176B-parameter open-access multilingual language model,” arXiv preprint, arXiv:2211.05100, 2023.

[23] Y. Bang et al., “A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity,” in Proc. IJCNLP-AACL, 2023, pp. 675–718.

[24] [K. Ahuja et al., “MEGA: Multilingual evaluation of generative AI,” in Proc. EMNLP, 2023, pp. 4232–4267.

[25] Anonymous, “Towards inclusive NLP: Evaluating LLMs on low-resource Indo-Iranian languages,” NeurIPS submission, 2025.

[26] J. Nay, “Natural language processing and machine learning for law and policy texts,” SSRN Electronic Journal, 2018, doi: 10.2139/ssrn.3438276.

[27] D. Premasiri et al., “Survey on legal information extraction: Current status and open challenges,” Knowledge and Information Systems, vol. 67, pp. 11287–11358, 2025, doi: 10.1007/s10115-025-02600-5.

[28] U. Khalid, “Natural language processing in legal document analysis,” Multidisciplinary Research in Computing Information Systems, vol. 4, no. 2, pp. 87–97, 2024.

[29] I. Trancoso et al., “The impact of language technologies in the legal domain,” in Multidisciplinary Perspectives on AI and the Law, Springer, 2024.

[30] E. Bertoni et al., Handbook of Computational Social Science for Policy. Springer, 2023.

[31] P. Roy et al., (if additional low-resource or ensemble study included separately; adjust numbering if needed).

How to cite this paper

Pranjal Gahlot, Prof. Rakshitha BS "A Multi-Metric Evaluation Perspective on Hallucination Detection in Low-Resource Governance Documents" Iconic Research And Engineering Journals Volume 9 Issue 11 2026 Page 2631-2642 https://doi.org/10.64388/IREV9I11-1717980
Pranjal Gahlot, Prof. Rakshitha BS "A Multi-Metric Evaluation Perspective on Hallucination Detection in Low-Resource Governance Documents" Iconic Research And Engineering Journals, vol. 9, no. 11, May. 2026, doi: https://doi.org/10.64388/IREV9I11-1717980
Pranjal Gahlot, Prof. Rakshitha BS (2026). A Multi-Metric Evaluation Perspective on Hallucination Detection in Low-Resource Governance Documents. Iconic Research And Engineering Journals, 9(11). doi: https://doi.org/10.64388/IREV9I11-1717980
Pranjal Gahlot, Prof. Rakshitha BS "A Multi-Metric Evaluation Perspective on Hallucination Detection in Low-Resource Governance Documents" Iconic Research And Engineering Journals, vol. 9, no. 11, May. 2026. Crossref, https://doi.org/10.64388/IREV9I11-1717980
@article{1717980,
      author = {Pranjal Gahlot, Prof. Rakshitha BS},
      title = {A Multi-Metric Evaluation Perspective on Hallucination Detection in Low-Resource Governance Documents},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {9},
      number = {11},
      pages = {2631-2642},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1717980.pdf},
      abstract = {The rapid advancement of Large Language Models (LLMs) has significantly improved natural language processing applications across domains such as governance, healthcare, legal analysis, and public information systems. Despite these advancements, LLMs frequently generate hallucinated outputs, where responses appear plausible but contain incorrect or fabricated information. This issue poses serious risks in governance-related applications, where inaccurate information can influence policy interpretation, administrative decision-making, and public trust. Existing studies have proposed several approaches to address hallucinations, including semantic entropy–based detection, benchmark evaluation frameworks, and adversarial testing methods. However, the literature indicates that current solutions remain fragmented and often focus on isolated aspects such as model performance, dataset construction, or benchmark capability rather than comprehensive reliability assessment. This literature review examines recent research on hallucination detection, multilingual and low-resource natural language processing, and evaluation frameworks for LLM reliability. The reviewed studies highlight key challenges, including the lack of multilingual hallucination evaluation, insufficient harm-oriented risk assessment, and limited adversarial robustness testing in governance contexts. Furthermore, existing benchmarks often measure task accuracy rather than factual reliability or societal impact. Based on the analysis of the literature, this review identifies major methodological and contextual gaps and proposes the need for an integrated evaluation framework combining meaning level hallucination detection, harm aware risk modeling, and multilingual robustness assessment. Such an approach could improve the reliability and safety of LLM systems deployed in governance and public service environments.},
      keywords = {Large Language Models, Hallucination Detection, Semantic Entropy, Multilingual NLP, Low-Resource Languages, Governance AI, Adversarial Prompting, Benchmark Evaluation, AI Reliability, Natural Language Processing.},
      month = {May},
      doi = {https://doi.org/10.64388/IREV9I11-1717980}
  }