International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1722808

1722808 Vol 8 · Issue 2 Download Paper

Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers

Srikumar Nayak Tanya Jane

Subject area: Science,Engineering and Technology  ·  Area of research: Detecting Prompt Injection

DOI: 10.64388/IREV8I2-1722808

Abstract

Retrieval-Augmented Generation (RAG) grounds large language model (LLM) outputs in externally retrieved documents, but this same mechanism creates a novel attack surface: adversarial instructions embedded inside retrieved content can hijack the downstream model's behavior, a threat known as indirect prompt injection. Unlike classical adversarial machine learning, which perturbs input features to evade a classifier, indirect prompt injection exploits the inability of an LLM to structurally distinguish trusted developer instructions from untrusted retrieved data. This paper proposes and evaluates a lightweight, model-agnostic defense: a supervised classifier that screens retrieved passages for injected instructions before they reach the generation model's context window. We construct a synthetic, multi-domain RAG corpus spanning banking, healthcare, life sciences, general technology, and customer-support retrieval scenarios, and train three lightweight classifiers — logistic regression, a calibrated linear support vector machine, and a random forest — on TF-IDF features. On a held-out test set the best classifier achieves near-perfect detection (F1 = 1.000, AUC = 1.000). Critically, we then evaluate generalization against an adaptive-obfuscation test set using homoglyph substitution, payload splitting, and unseen synonym paraphrasing not present in training; recall degrades to 89.5–92.0% while precision remains at 100%, revealing an asymmetric robustness gap between naive and adaptive attackers. We further quantify the false-positive-rate/usability tradeoff via a decision-threshold sweep, showing how operators can tune the detector for high-recall security-critical deployments (e.g., banking transaction assistants) versus low-friction deployments (e.g., customer support). We discuss deployment implications for banking, healthcare, and life-sciences RAG systems, where both missed injections and false blocks carry asymmetric operational costs, and outline future work on adversarially robust training and cross-lingual obfuscation.

Keywords

Prompt injection, retrieval-augmented generation, large language models, adversarial machine learning, intrusion detection, TF-IDF, lightweight classifiers, LLM security, AI safety, banking security, healthcare AI, life sciences AI.

References

[1] P. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Proc. NeurIPS, 2020, pp. 9459-9474.

[2] F. Perez and I. Ribeiro, "Ignore previous prompt: Attack techniques for language models," arXiv:2211.09527, 2022.

[3] K. Greshake et al., "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection," arXiv:2302.12173, 2023.

[4] I. J. Goodfellow, J. Shlens, and C. Szegedy, "Explaining and harnessing adversarial examples," in Proc. ICLR, 2015.

[5] A. Vaswani et al., "Attention is all you need," in Proc. NeurIPS, 2017, pp. 5998-6008.

[6] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. NAACL-HLT, 2019, pp. 4171-4186.

[7] T. Wolf et al., "Transformers: State-of-the-art natural language processing," in Proc. EMNLP: Syst. Demonstrations, 2020, pp. 38-45.

[8] F. T. Liu, K. M. Ting, and Z.-H. Zhou, "Isolation forest," in Proc. IEEE ICDM, 2008, pp. 413-422.

[9] V. Chandola, A. Banerjee, and V. Kumar, "Anomaly detection: A survey," ACM Comput. Surv., vol. 41, no. 3, pp. 1-58, 2009.

[10] O. K. Sahingoz, E. Buber, O. Demir, and B. Diri, "Machine learning based phishing detection from URLs," Expert Syst. Appl., vol. 117, pp. 345-357, 2019.

[11] F. Pedregosa et al., "Scikit-learn: Machine learning in Python," J. Mach. Learn. Res., vol. 12, pp. 2825-2830, 2011.

[12] R. Sommer and V. Paxson, "Outside the closed world: On using machine learning for network intrusion detection," in Proc. IEEE Symp. Security and Privacy, 2010, pp. 305-316.

[13] M. Roesch, "Snort: Lightweight intrusion detection for networks," in Proc. USENIX LISA, 1999, pp. 229-238.

[14] K. Rieck and P. Laskov, "Language models for detection of unknown attacks in network traffic," J. Comput. Virology, vol. 2, no. 4, pp. 243-256, 2007.

[15] A. L. Buczak and E. Guven, "A survey of data mining and machine learning methods for cyber security intrusion detection," IEEE Commun. Surveys Tuts., vol. 18, no. 2, pp. 1153-1176, 2016.

[16] G. Apruzzese, M. Colajanni, L. Ferretti, A. Guido, and M. Marchetti, "On the effectiveness of machine and deep learning for cyber security," in Proc. Int. Conf. Cyber Conflict (CyCon), 2018, pp. 371-390.

[17] K. Rieck, P. Trinius, C. Willems, and T. Holz, "Automatic analysis of malware behavior using machine learning," J. Comput. Security, vol. 19, no. 4, pp. 639-668, 2011.

[18] J. Saxe and K. Berlin, "Deep neural network based malware detection using two dimensional binary program features," in Proc. Int. Conf. Malicious and Unwanted Software (MALWARE), 2015, pp. 11-20.

[19] C. Szegedy et al., "Intriguing properties of neural networks," in Proc. ICLR, 2014.

[20] N. Papernot, P. McDaniel, S. Jha, M. Fredrikson, Z. B. Celik, and A. Swami, "The limitations of deep learning in adversarial settings," in Proc. IEEE Eur. Symp. Security and Privacy, 2016, pp. 372-387.

[21] N. Carlini and D. Wagner, "Towards evaluating the robustness of neural networks," in Proc. IEEE Symp. Security and Privacy, 2017, pp. 39-57.

[22] T. Gu, B. Dolan-Gavitt, and S. Garg, "BadNets: Identifying vulnerabilities in the machine learning model supply chain," arXiv:1708.06733, 2017.

[23] X. Chen, C. Liu, B. Li, K. Lu, and D. Song, "Targeted backdoor attacks on deep learning systems using data poisoning," arXiv:1712.05526, 2017.

[24] F. Tramer, F. Zhang, A. Juels, M. K. Reiter, and T. Ristenpart, "Stealing machine learning models via prediction APIs," in Proc. USENIX Security Symp., 2016, pp. 601-618.

[25] M. Fredrikson, S. Jha, and T. Ristenpart, "Model inversion attacks that exploit confidence information and basic countermeasures," in Proc. ACM SIGSAC CCS, 2015, pp. 1322-1333.

[26] R. Shokri, M. Stronati, C. Song, and V. Shmatikov, "Membership inference attacks against machine learning models," in Proc. IEEE Symp. Security and Privacy, 2017, pp. 3-18.

[27] B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, "Communication-efficient learning of deep networks from decentralized data," in Proc. AISTATS, 2017, pp. 1273-1282.

[28] P. Kairouz et al., "Advances and open problems in federated learning," Foundations and Trends in Machine Learning, vol. 14, no. 1-2, pp. 1-210, 2021.

[29] M. T. Ribeiro, S. Singh, and C. Guestrin, "'Why should I trust you?': Explaining the predictions of any classifier," in Proc. ACM SIGKDD, 2016, pp. 1135-1144.

[30] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Proc. NeurIPS, 2017, pp. 4765-4774.

[31] I. Goodfellow et al., "Generative adversarial networks," in Proc. NeurIPS, 2014, pp. 2672-2680.

[32] A. Rossler, D. Cozzolino, L. Verdoliva, C. Riess, J. Thies, and M. Niessner, "FaceForensics++: Learning to detect manipulated facial images," in Proc. IEEE/CVF ICCV, 2019, pp. 1-11.

[33] G. Salton and C. Buckley, "Term-weighting approaches in automatic text retrieval," Inf. Process. Manage., vol. 24, no. 5, pp. 513-523, 1988.

[34] T. Joachims, "Text categorization with support vector machines: Learning with many relevant features," in Proc. ECML, 1998, pp. 137-142.

[35] L. Breiman, "Random forests," Machine Learning, vol. 45, no. 1, pp. 5-32, 2001.

[36] C. Cortes and V. Vapnik, "Support-vector networks," Machine Learning, vol. 20, no. 3, pp. 273-297, 1995.

How to cite this paper

Srikumar Nayak, Tanya Jane "Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers" Iconic Research And Engineering Journals Volume 8 Issue 2 2024 Page 1428-1437 https://doi.org/10.64388/IREV8I2-1722808
Srikumar Nayak, Tanya Jane "Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers" Iconic Research And Engineering Journals, vol. 8, no. 2, Aug. 2024, doi: https://doi.org/10.64388/IREV8I2-1722808
Srikumar Nayak, Tanya Jane (2024). Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers. Iconic Research And Engineering Journals, 8(2). doi: https://doi.org/10.64388/IREV8I2-1722808
Srikumar Nayak, Tanya Jane "Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers" Iconic Research And Engineering Journals, vol. 8, no. 2, Aug. 2024. Crossref, https://doi.org/10.64388/IREV8I2-1722808
@article{1722808,
      author = {Srikumar Nayak, Tanya Jane},
      title = {Detecting Prompt Injection in Retrieval-Augmented Generation Pipelines Using Lightweight Classifiers},
      journal = {Iconic Research And Engineering Journals},
      year = {2024},
      volume = {8},
      number = {2},
      pages = {1428-1437},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1722808.pdf},
      abstract = {Retrieval-Augmented Generation (RAG) grounds large language model (LLM) outputs in externally retrieved documents, but this same mechanism creates a novel attack surface: adversarial instructions embedded inside retrieved content can hijack the downstream model's behavior, a threat known as indirect prompt injection. Unlike classical adversarial machine learning, which perturbs input features to evade a classifier, indirect prompt injection exploits the inability of an LLM to structurally distinguish trusted developer instructions from untrusted retrieved data. This paper proposes and evaluates a lightweight, model-agnostic defense: a supervised classifier that screens retrieved passages for injected instructions before they reach the generation model's context window. We construct a synthetic, multi-domain RAG corpus spanning banking, healthcare, life sciences, general technology, and customer-support retrieval scenarios, and train three lightweight classifiers — logistic regression, a calibrated linear support vector machine, and a random forest — on TF-IDF features. On a held-out test set the best classifier achieves near-perfect detection (F1 = 1.000, AUC = 1.000). Critically, we then evaluate generalization against an adaptive-obfuscation test set using homoglyph substitution, payload splitting, and unseen synonym paraphrasing not present in training; recall degrades to 89.5–92.0% while precision remains at 100%, revealing an asymmetric robustness gap between naive and adaptive attackers. We further quantify the false-positive-rate/usability tradeoff via a decision-threshold sweep, showing how operators can tune the detector for high-recall security-critical deployments (e.g., banking transaction assistants) versus low-friction deployments (e.g., customer support). We discuss deployment implications for banking, healthcare, and life-sciences RAG systems, where both missed injections and false blocks carry asymmetric operational costs, and outline future work on adversarially robust training and cross-lingual obfuscation.},
      keywords = {Prompt injection, retrieval-augmented generation, large language models, adversarial machine learning, intrusion detection, TF-IDF, lightweight classifiers, LLM security, AI safety, banking security, healthcare AI, life sciences AI.},
      month = {August},
      doi = {https://doi.org/10.64388/IREV8I2-1722808}
  }