Imperceptible Content Poisoning in LLM-Powered Applications
Quan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng, Heyuan Shi, Zichen Xu, Yu Jiang
Abstract
Large Language Models (LLMs) have shown their superior capability in natural language processing, promoting extensive LLM-powered applications to be the new portals for people to access various content on the Internet. However, LLM-powered applications do not have sufficient security considerations on untrusted content, leading to potential threats. In this paper, we reveal content poisoning, where attackers can tailor attack content that appears benign to humans but causes LLM-powered applications to generate malicious responses. To highlight the impact of content poisoning and inspire the development of effective defenses, we systematically analyze the attack, focusing on the attack modes in various content, exploitable design features of LLM application frameworks, and the generation of attack content. We carry out a comprehensive evaluation on five LLMs, where content poisoning achieves an average attack success rate of 89.60%. Additionally, we assess content poisoning on four popular LLM-powered applications, achieving the attack on 72.00% of the content. Our experimental results also show that existing defenses are ineffective against content poisoning. Finally, we discuss potential mitigations for LLM application frameworks to counter content poisoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM AgentsHaoyu Wang, Christopher M. Poskitt, Jun SunICSE 2026 · 2 citations
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering AttractorsYen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung ChenICML 2026
Builds on8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li et al.ICLR 2024 · 481 citations
- AdvDoor: adversarial backdoor attack of deep learning systemQuan Zhang, Yifeng Ding, Yongqiang Tian, Jianmin Guo et al.ISSTA 2021 · 57 citations
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman et al.KDD 2025 · 27 citations
- MAWSEO: Adversarial Wiki Search Poisoning for Illicit Online PromotionZilong Lin, Zhengyi Li, Xiaojing Liao, XiaoFeng Wang et al.S&P 2024 · 16 citations
Related papers
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen et al.ICML 2025
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang et al.NDSS 2026 · 6 citations
- Are LLM-Enhanced Graph Neural Networks Robust Against Poisoning Attacks?Yuhang Ma, Jie Wang, Zheng YanS&P 2026 · 4 citations
- Prompt-to-SQL Injections in LLM-Integrated Web Applications: Risks and DefensesRodrigo Pedro, Miguel E. Coimbra, Daniel Castro, Paulo Carreira et al.ICSE 2025 · 14 citations
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina et al.CCS 2024 · 28 citations
