Home / Current Issue / Paper 1722808
Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers
Subject area: Science,Engineering and Technology · Area of research: Detecting Prompt Injection
Abstract
Retrieval-Augmented Generation (RAG) grounds large language model (LLM) outputs in externally retrieved documents, but this same mechanism creates a novel attack surface: adversarial instructions embedded inside retrieved content can hijack the downstream model's behavior, a threat known as indirect prompt injection. Unlike classical adversarial machine learning, which perturbs input features to evade a classifier, indirect prompt injection exploits the inability of an LLM to structurally distinguish trusted developer instructions from untrusted retrieved data. This paper proposes and evaluates a lightweight, model-agnostic defense: a supervised classifier that screens retrieved passages for injected instructions before they reach the generation model's context window. We construct a synthetic, multi-domain RAG corpus spanning banking, healthcare, life sciences, general technology, and customer-support retrieval scenarios, and train three lightweight classifiers — logistic regression, a calibrated linear support vector machine, and a random forest — on TF-IDF features. On a held-out test set the best classifier achieves near-perfect detection (F1 = 1.000, AUC = 1.000). Critically, we then evaluate generalization against an adaptive-obfuscation test set using homoglyph substitution, payload splitting, and unseen synonym paraphrasing not present in training; recall degrades to 89.5–92.0% while precision remains at 100%, revealing an asymmetric robustness gap between naive and adaptive attackers. We further quantify the false-positive-rate/usability tradeoff via a decision-threshold sweep, showing how operators can tune the detector for high-recall security-critical deployments (e.g., banking transaction assistants) versus low-friction deployments (e.g., customer support). We discuss deployment implications for banking, healthcare, and life-sciences RAG systems, where both missed injections and false blocks carry asymmetric operational costs, and outline future work on adversarially robust training and cross-lingual obfuscation.
Keywords
Prompt injection, retrieval-augmented generation, large language models, adversarial machine learning, intrusion detection, TF-IDF, lightweight classifiers, LLM security, AI safety, banking security, healthcare AI, life sciences AI.
References
[1] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. NeurIPS, 2020, pp. 9459-9474.
[2] F. Perez and I. Ribeiro, "Ignore previous prompt: Attack techniques for language models," arXiv:2211.09527, 2022.
[3] K. Greshake et al., "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection," arXiv:2302.12173, 2023.
[4] I. J. Goodfellow, J. Shlens, and C. Szegedy, "Explaining and harnessing adversarial examples," in Proc. ICLR, 2015.
[5] A. Vaswani et al., "Attention is all you need," in Proc. NeurIPS, 2017, pp. 5998-6008.
[6] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT, 2019, pp. 4171-4186.
[7] T. Wolf et al., "Transformers: State-of-the-art natural language processing," in Proc. EMNLP: Syst. Demonstrations, 2020, pp. 38-45.
[8] F. T. Liu, K. M. Ting, and Z.-H. Zhou, "Isolation forest," in Proc. IEEE ICDM, 2008, pp. 413-422.
[9] V. Chandola, A. Banerjee, and V. Kumar, "Anomaly detection: A survey," ACM Comput. Surv., vol. 41, no. 3, pp. 1-58, 2009.
[10] O. K. Sahingoz, E. Buber, O. Demir, and B. Diri, "Machine learning based phishing detection from URLs," Expert Syst. Appl., vol. 117, pp. 345-357, 2019.
[11] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825-2830, 2011.
[12] R. Sommer and V. Paxson, "Outside the closed world: On using machine learning for network intrusion detection," in Proc. IEEE Symp. Security and Privacy, 2010, pp. 305-316.
[13] M. Roesch, "Snort: Lightweight intrusion detection for networks," in Proc. USENIX LISA, 1999, pp. 229-238.
[14] K. Rieck and P. Laskov, "Language models for detection of unknown attacks in network traffic," J. Comput. Virology, vol. 2, no. 4, pp. 243-256, 2007.
[15] A. L. Buczak and E. Guven, "A survey of data mining and machine learning methods for cyber security intrusion detection," IEEE Commun. Surveys Tuts., vol. 18, no. 2, pp. 1153-1176, 2016.
[16] G. Apruzzese, M. Colajanni, L. Ferretti, A. Guido, and M. Marchetti, "On the effectiveness of machine and deep learning for cyber security," in Proc. Int. Conf. Cyber Conflict (CyCon), 2018, pp. 371-390.
[17] K. Rieck, P. Trinius, C. Willems, and T. Holz, "Automatic analysis of malware behavior using machine learning," J. Comput. Security, vol. 19, no. 4, pp. 639-668, 2011.
[18] J. Saxe and K. Berlin, "Deep neural network based malware detection using two dimensional binary program features," in Proc. Int. Conf. Malicious and Unwanted Software (MALWARE), 2015, pp. 11-20.
[19] C. Szegedy et al., "Intriguing properties of neural networks," in Proc. ICLR, 2014.
[20] N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z. B. Celik, and A. Swami, "The limitations of deep learning in adversarial settings," in Proc. IEEE Eur. Symp. Security and Privacy, 2016, pp. 372-387.
[21] N. Carlini and D. Wagner, "Towards evaluating the robustness of neural networks," in Proc. IEEE Symp. Security and Privacy, 2017, pp. 39-57.
[22] T. Gu, B. Dolan-Gavitt, and S. Garg, "BadNets: Identifying vulnerabilities in the machine learning model supply chain," arXiv:1708.06733, 2017.
[23] X. Chen, C. Liu, B. Li, K. Lu, and D. Song, "Targeted backdoor attacks on deep learning systems using data poisoning," arXiv:1712.05526, 2017.
[24] F. Tramer, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart, "Stealing machine learning models via prediction APIs," in Proc. USENIX Security Symp., 2016, pp. 601-618.
[25] M. Fredrikson, S. Jha, and T. Ristenpart, "Model inversion attacks that exploit confidence information and basic countermeasures," in Proc. ACM SIGSAC CCS, 2015, pp. 1322-1333.
[26] R. Shokri, M. Stronati, C. Song, and V. Shmatikov, "Membership inference attacks against machine learning models," in Proc. IEEE Symp. Security and Privacy, 2017, pp. 3-18.
[27] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, "Communication-efficient learning of deep networks from decentralized data," in Proc. AISTATS, 2017, pp. 1273-1282.
[28] P. Kairouz et al., "Advances and open problems in federated learning," Foundations and Trends in Machine Learning, vol. 14, no. 1-2, pp. 1-210, 2021.
[29] M. T. Ribeiro, S. Singh, and C. Guestrin, "'Why should I trust you?': Explaining the predictions of any classifier," in Proc. ACM SIGKDD, 2016, pp. 1135-1144.
[30] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Proc. NeurIPS, 2017, pp. 4765-4774.
[31] I. Goodfellow et al., "Generative adversarial networks," in Proc. NeurIPS, 2014, pp. 2672-2680.
[32] A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Niessner, "FaceForensics++: Learning to detect manipulated facial images," in Proc. IEEE/CVF ICCV, 2019, pp. 1-11.
[33] G. Salton and C. Buckley, "Term-weighting approaches in automatic text retrieval," Inf. Process. Manage., vol. 24, no. 5, pp. 513-523, 1988.
[34] T. Joachims, "Text categorization with support vector machines: Learning with many relevant features," in Proc. ECML, 1998, pp. 137-142.
[35] L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5-32, 2001.
[36] C. Cortes and V. Vapnik, "Support-vector networks," Machine Learning, vol. 20, no. 3, pp. 273-297, 1995.
How to cite this paper
@article{1722808,
author = {Srikumar Nayak, Tanya Jane},
title = {Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers},
journal = {Iconic Research And Engineering Journals},
year = {2024},
volume = {8},
number = {2},
pages = {1428-1437},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1722808.pdf},
abstract = {Retrieval-Augmented Generation (RAG) grounds large language model (LLM) outputs in externally retrieved documents, but this same mechanism creates a novel attack surface: adversarial instructions embedded inside retrieved content can hijack the downstream model's behavior, a threat known as indirect prompt injection. Unlike classical adversarial machine learning, which perturbs input features to evade a classifier, indirect prompt injection exploits the inability of an LLM to structurally distinguish trusted developer instructions from untrusted retrieved data. This paper proposes and evaluates a lightweight, model-agnostic defense: a supervised classifier that screens retrieved passages for injected instructions before they reach the generation model's context window. We construct a synthetic, multi-domain RAG corpus spanning banking, healthcare, life sciences, general technology, and customer-support retrieval scenarios, and train three lightweight classifiers — logistic regression, a calibrated linear support vector machine, and a random forest — on TF-IDF features. On a held-out test set the best classifier achieves near-perfect detection (F1 = 1.000, AUC = 1.000). Critically, we then evaluate generalization against an adaptive-obfuscation test set using homoglyph substitution, payload splitting, and unseen synonym paraphrasing not present in training; recall degrades to 89.5–92.0% while precision remains at 100%, revealing an asymmetric robustness gap between naive and adaptive attackers. We further quantify the false-positive-rate/usability tradeoff via a decision-threshold sweep, showing how operators can tune the detector for high-recall security-critical deployments (e.g., banking transaction assistants) versus low-friction deployments (e.g., customer support). We discuss deployment implications for banking, healthcare, and life-sciences RAG systems, where both missed injections and false blocks carry asymmetric operational costs, and outline future work on adversarially robust training and cross-lingual obfuscation.},
keywords = {Prompt injection, retrieval-augmented generation, large language models, adversarial machine learning, intrusion detection, TF-IDF, lightweight classifiers, LLM security, AI safety, banking security, healthcare AI, life sciences AI.},
month = {August},
doi = {https://doi.org/10.64388/IREV8I2-1722808}
}