Can Indirect Prompt Injection Attacks Be Detected and Removed?
Yulin Chen, Haoran Li, Yuan Sui, Yufei He, Yue Liu, Yangqiu Song, Bryan Hooi
Abstract
Prompt injection attacks manipulate large language models (LLMs) by misleading them to deviate from the original input instructions and execute maliciously injected instructions, because of their instruction-following capabilities and inability to distinguish between the original input instructions and maliciously injected instructions. To defend against such attacks, recent studies have developed various detection mechanisms. If we restrict ourselves specifically to works which perform detection rather than direct defense, most of them focus on direct prompt injection attacks, while there are few works for the indirect scenario, where injected instructions are indirectly from external tools, such as a search engine. Moreover, current works mainly investigate injection detection methods and pay less attention to the post-processing method that aims to mitigate the injection after detection. In this paper, we investigate the feasibility of detecting and removing indirect prompt injection attacks, and we construct a benchmark dataset for evaluation. For detection, we assess the performance of existing LLMs and open-source detection models, and we further train detection models using our crafted training datasets. For removal, we evaluate two intuitive methods: (1) the segmentation removal method, which segments the injected document and removes parts containing injected instructions, and (2) the extraction removal method, which trains an extraction model to identify and remove injected instructions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0b1dd26e-d79c-4204-99a6-7d86ed71207dCited by top-tier papers9
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic SystemsYufei He, Juncheng Liu, Yue Liu, Yibo Li et al.ICLR 2026 · 36 citations
- PromptLocate: Localizing Prompt Injection AttacksYuqi Jia, Yupei Liu, Zedian Shao, Jinyuan Jia et al.S&P 2026 · 35 citations
- ASIDE: Architectural Separation of Instructions and Data in Language ModelsEgor Zverev, Evgenii Kortukov, Alexander Panfilov, Alexandra Volkova et al.ICLR 2026 · 28 citations
- VIGIL: Defending LLM Agents Against Tool-Stream Injection via Verify-Before-CommitJunda Lin, Zhaomeng Zhou, Zhi Zheng, Shuochen Liu et al.ACL 2026 · 7 citations
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following BehaviorWeikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li et al.CVPR 2026 · 6 citations
Builds on11
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Least-to-Most Prompting Enables Complex Reasoning in Large Language ModelsDenny Zhou, Nathanael Schärli, Le Hou, Jason Wei et al.ICLR 2023 · 318 citations
- On the Exploitability of Instruction TuningManli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping et al.NeurIPS 2023 · 166 citations
- KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing DetectionYuexin Li, Chengyu Huang, Shumin Deng, Mei Lin Lock et al.USENIX Security 2024 · 70 citations
Related papers
- TopicAttack: An Indirect Prompt Injection Attack via Topic TransitionYulin Chen, Haoran Li, Yuexin Li, Yue Liu et al.EMNLP 2025 · 1 citation
- Tensor Trust: Interpretable Prompt Injection Attacks from an Online GameSam Toyer, Olivia Watkins, Ethan Adrian Mendes, Justin Svegliato et al.ICLR 2024 · 123 citations
- Evaluating the Instruction-Following Robustness of Large Language Models to Prompt InjectionZekun Li, Baolin Peng, Pengcheng He, Xifeng YanEMNLP 2024 · 15 citations
- Defense Against Prompt Injection Attack by Leveraging Attack TechniquesYulin Chen, Haoran Li, Zihao Zheng, Dekai Wu et al.ACL 2025
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
