International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1723870

1723870 Vol 10 · Issue 4 Download Paper

Retrieval Quality, Evidence, and Appropriate Abstention in Large Language Models

Purvish Haresh Sharma

Subject area: Science,Engineering and Technology  ·  Area of research: Artificial Intelligence & Machine Learning

Abstract

Large Language Models (LLMs) can give fluent and useful answers, but they can also produce information that is false or not supported by evidence. This problem is often called hallucination. In this paper, we study how the quality of the evidence retrieved is associated with differences in the answers given by an LLM. The experiment used 13 questions in English, Hindi, and Gujarati. The questions were tested under three conditions: No Retrieval, High-quality Retrieval, and Low-quality Retrieval. In total, the experiment produced 117 responses. The responses were evaluated using factual correctness, evidence support, appropriate abstention, over-cautious behavior, unsupported inference, and non-evaluable cases. In the No Retrieval condition, all 39 responses were fully correct. In the High-quality Retrieval condition, 34 of 36 evaluable responses were fully correct. In the Low-quality Retrieval condition, only 3 of 39 responses were fully correct, but 36 responses showed appropriate abstention. This means that the model often responded that the evidence was not sufficient instead of making an unsupported claim. The results suggest that retrieval does not automatically improve the reliability of an LLM. The quality, relevance and completeness of the evidence are also important. The study also shows that factual correctness, evidence faithfulness, and appropriate uncertainty should be evaluated separately.

Keywords

large language models; hallucination; retrieval-augmented generation; evidence faithfulness; abstention; multilingual evaluation

References

[1] Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of hallucination in natural language generation. ACM Computing Surveys. 2023;55(12):248. https://doi.org/10.1145/3571730

[2] Huang L, Yu W, Ma W, Zhong W, Feng Z, Wang H, et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems. 2025;43(2):42. https://doi.org/10.1145/3703155

[3] Wang Y, Wang M, Manzoor MA, Liu F, Georgiev GN, Das RJ, et al. Factuality of large language models: a survey. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024); 2024. p. 19519–19529. https://doi.org/10.18653/v1/2024.emnlp-main.1088

[4] Venkit PN, Chakravorti T, Gupta V, Biggs H, Srinath M, Goswami K, et al. An audit on the perspectives and challenges of hallucinations in NLP. In: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP 2024); 2024. p. 6528–6548. https://doi.org/10.18653/v1/2024.emnlp-main.375

[5] Li J, Chen J, Ren R, Cheng X, Zhao X, Nie JY, et al. The dawn after the dark: an empirical study on factuality hallucination in large language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024); 2024. p. 10879–10899. https://doi.org/10.18653/v1/2024.acl-long.586

[6] Tonmoy SMTI, Zaman SMM, Jain V, Rani A, Rawte V, Chadha A, et al. A comprehensive survey of hallucination mitigation techniques in large language models [preprint]. arXiv:2401.01313; 2024. https://doi.org/10.48550/arXiv.2401.01313

[7] Pan L, Saxon M, Xu W, Nathani D, Wang X, Wang WY. Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies. Transactions of the Association for Computational Linguistics. 2024;12:484–506. https://doi.org/10.1162/tacl_a_00660

[8] Sahoo P, Meharia P, Ghosh A, Saha S, Jain V, Chadha A. A comprehensive survey of hallucination in large language, image, video and audio foundation models. In: Findings of the Association for Computational Linguistics: EMNLP 2024; 2024. p. 11709–11724. https://doi.org/10.18653/v1/2024.findings-emnlp.685

[9] van Deemter K. The pitfalls of defining hallucination. Computational Linguistics. 2024;50(2):807–816. https://doi.org/10.1162/coli_a_00509

[10] Saxena A, Bhattacharyya P. Hallucination detection in machine generated text: a survey. Mumbai: Department of Computer Science and Engineering, Indian Institute of Technology Bombay; 2024.

[11] Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In: Advances in Neural Information Processing Systems 33 (NeurIPS 2020); 2020. p. 9459–9474. https://doi.org/10.48550/arXiv.2005.11401

[12] Niu C, Wu Y, Zhu J, Xu S, Shum K, Zhong R, et al. RAGTruth: a hallucination corpus for developing trustworthy retrieval-augmented language models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024); 2024. p. 10862–10878. https://doi.org/10.18653/v1/2024.acl-long.585

[13] Google. Gemini 3.5 Flash-Lite [Internet]. Google AI for Developers; 2026. Available from: https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash-lite

How to cite this paper

Purvish Haresh Sharma "Retrieval Quality, Evidence, and Appropriate Abstention in Large Language Models" Iconic Research And Engineering Journals Volume 10 Issue 4 2026 Page 1520-1536
Purvish Haresh Sharma "Retrieval Quality, Evidence, and Appropriate Abstention in Large Language Models" Iconic Research And Engineering Journals, vol. 10, no. 4, Oct. 2026
Purvish Haresh Sharma (2026). Retrieval Quality, Evidence, and Appropriate Abstention in Large Language Models. Iconic Research And Engineering Journals, 10(4).
Purvish Haresh Sharma "Retrieval Quality, Evidence, and Appropriate Abstention in Large Language Models" Iconic Research And Engineering Journals, vol. 10, no. 4, Oct. 2026.
@article{1723870,
      author = {Purvish Haresh Sharma},
      title = {Retrieval Quality, Evidence, and Appropriate Abstention in Large Language Models},
      journal = {Iconic Research And Engineering Journals},
      year = {2026},
      volume = {10},
      number = {4},
      pages = {1520-1536},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1723870.pdf},
      abstract = {Large Language Models (LLMs) can give fluent and useful answers, but they can also produce information that is false or not supported by evidence. This problem is often called hallucination. In this paper, we study how the quality of the evidence retrieved is associated with differences in the answers given by an LLM.
The experiment used 13 questions in English, Hindi, and Gujarati. The questions were tested under three conditions: No Retrieval, High-quality Retrieval, and Low-quality Retrieval. In total, the experiment produced 117 responses. The responses were evaluated using factual correctness, evidence support, appropriate abstention, over-cautious behavior, unsupported inference, and non-evaluable cases.
In the No Retrieval condition, all 39 responses were fully correct. In the High-quality Retrieval condition, 34 of 36 evaluable responses were fully correct. In the Low-quality Retrieval condition, only 3 of 39 responses were fully correct, but 36 responses showed appropriate abstention. This means that the model often responded that the evidence was not sufficient instead of making an unsupported claim.
The results suggest that retrieval does not automatically improve the reliability of an LLM. The quality, relevance and completeness of the evidence are also important. The study also shows that factual correctness, evidence faithfulness, and appropriate uncertainty should be evaluated separately.},
      keywords = {large language models; hallucination; retrieval-augmented generation; evidence faithfulness; abstention; multilingual evaluation},
      month = {October},
  }