Home / Current Issue / Paper 1712249
Retrieval-Augmented Generation (RAG)-Based Chatbot System
Subject area: Science,Engineering and Technology · Area of research: Generative AI
Abstract
The fast development of Large Language Models (LLMs) has highly improved abilities in natural language processing, but deploying LLMs in high-stakes contexts continues to be problematic because of inaccuracy and hallucinations. In this paper, we describe a Retrieval-Augmented Generation (RAG) based chatbot system that avoids these issues from existing models and ensures all responses are grounded and based on valid knowledge sources. The RAG system builds an architecture that scans specific documents in a domain, derives semantic embeddings that include vectors, and stores their vectors in a FAISS vector database, combining retrieval while ensuring speed and memory efficiency. When a question is passed to the RAG bot, the bot retrieves relevant context passages and contexts to condition a generative language model to yield an accurate answer that is also cited. Our implementation involves a modular pipeline and allows all knowledge to be fresh, without having to retrain the language model. We provide experimental results showing significant improvements in factual accuracy compared to baseline LLMs and improved reductions in hallucination. Conclusion: the system is constructed in a domain-free manner allowing further employment in healthcare, legal, or enterprise settings, where it is important to provide cited and verifiable information. The development of this chatbot-status represents an important milestone toward creating trustworthy Artificial Intelligence (AI) systems that exhibit generative fluency and factual reliability.
Keywords
Retrieval-Augmented Generation, Large Language Models, Semantic Search, FAISS, Chatbot Systems, Hallucination Mitigation, Knowledge Grounding.
References
[1] A. Vaswani et al., "Attention is all you need," in Advances in Neural Information Processing Systems, 2017, pp. 5998-6008.
[2] J. Devlin et al., "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT, 2019, pp. 4171-4186.
[3] J. Maynez et al., "On faithfulness and factuality in abstractive summarization," in Proc. ACL, 2020, pp. 1906-1919.
[4] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. NeurIPS, 2020, pp. 9459-9474.
[5] R. McDonald et al., "WikiReading: A novel large-scale language understanding task over Wikipedia," in Proc. ACL, 2018.
[6] M. Lewis et al., "BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension," in Proc. ACL, 2020, pp. 7871-7880.
[7] P. Rajpurkar et al., "SQuAD: 100,000+ questions for machine comprehension of text," in Proc. EMNLP, 2016, pp. 2383-2392.
[8] V. Karpukhin et al., "Dense passage retrieval for open-domain question answering," in Proc. EMNLP, 2020, pp. 6769-6781.
[9] M. Lewis et al., "BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension," in Proc. ACL, 2020, pp. 7871-7880.
[10] K. Guu et al., "REALM: Retrieval-augmented language model pre-training," in Proc. ICML, 2020, pp. 345-356.
[11] "LangChain Documentation," [Online]. Available: https://docs.langchain.com/
[12] A. Johnson et al., "Medical question answering with retrieval-augmented generation," in Journal of Medical Systems, 2023.
[13] S. Chen et al., "Legal document analysis using hybrid retrieval-generation models," in Proc. ICAIL, 2023.
[14] M. Thompson et al., "Enterprise knowledge management with RAG systems," in Proc. KDD, 2023.
[15] N. Reimers and I. Gurevych, "Sentence-BERT: Sentence embeddings using Siamese BERT-networks," in Proc. EMNLP-IJCNLP, 2019, pp. 3982-3992.
[16] M. Douze et al., "Faiss: A library for efficient similarity search," Journal of Machine Learning Research, vol. 24, no. 1, pp. 1-6, 2023.
How to cite this paper
@article{1712249,
author = {Himank Garg, Anish Kumar, Himanshu Kumar, Ishrat Ali, Anuj Chandila},
title = {Retrieval-Augmented Generation (RAG)-Based Chatbot System},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {5},
pages = {1572-1575},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1712249.pdf},
abstract = {The fast development of Large Language Models (LLMs) has highly improved abilities in natural language processing, but deploying LLMs in high-stakes contexts continues to be problematic because of inaccuracy and hallucinations. In this paper, we describe a Retrieval-Augmented Generation (RAG) based chatbot system that avoids these issues from existing models and ensures all responses are grounded and based on valid knowledge sources. The RAG system builds an architecture that scans specific documents in a domain, derives semantic embeddings that include vectors, and stores their vectors in a FAISS vector database, combining retrieval while ensuring speed and memory efficiency. When a question is passed to the RAG bot, the bot retrieves relevant context passages and contexts to condition a generative language model to yield an accurate answer that is also cited. Our implementation involves a modular pipeline and allows all knowledge to be fresh, without having to retrain the language model. We provide experimental results showing significant improvements in factual accuracy compared to baseline LLMs and improved reductions in hallucination. Conclusion: the system is constructed in a domain-free manner allowing further employment in healthcare, legal, or enterprise settings, where it is important to provide cited and verifiable information. The development of this chatbot-status represents an important milestone toward creating trustworthy Artificial Intelligence (AI) systems that exhibit generative fluency and factual reliability.},
keywords = {Retrieval-Augmented Generation, Large Language Models, Semantic Search, FAISS, Chatbot Systems, Hallucination Mitigation, Knowledge Grounding.},
month = {November},
doi = {https://doi.org/10.64388/IREV9I5-1712249}
}