Imperceptible Content Poisoning in LLM-Powered Applications
Quan Zhang, Chijin Zhou, Gwihwan Go, Binqi Zeng, Heyuan Shi, Zichen Xu, Yu Jiang
摘要
Large Language Models (LLMs) have shown their superior capability in natural language processing, promoting extensive LLM-powered applications to be the new portals for people to access various content on the Internet. However, LLM-powered applications do not have sufficient security considerations on untrusted content, leading to potential threats. In this paper, we reveal content poisoning, where attackers can tailor attack content that appears benign to humans but causes LLM-powered applications to generate malicious responses. To highlight the impact of content poisoning and inspire the development of effective defenses, we systematically analyze the attack, focusing on the attack modes in various content, exploitable design features of LLM application frameworks, and the generation of attack content. We carry out a comprehensive evaluation on five LLMs, where content poisoning achieves an average attack success rate of 89.60%. Additionally, we assess content poisoning on four popular LLM-powered applications, achieving the attack on 72.00% of the content. Our experimental results also show that existing defenses are ineffective against content poisoning. Finally, we discuss potential mitigations for LLM application frameworks to counter content poisoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM AgentsHaoyu Wang, Christopher M. Poskitt, Jun SunICSE 2026 · 被引用 2 次
- Eyes-on-Me: Scalable RAG Poisoning through Transferable Attention-Steering AttractorsYen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung ChenICML 2026
它引用的顶会 Paper8
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Catastrophic Jailbreak of Open-source LLMs via Exploiting GenerationYangsibo Huang, Samyak Gupta, Mengzhou Xia, Kai Li 等ICLR 2024 · 被引用 481 次
- AdvDoor: adversarial backdoor attack of deep learning systemQuan Zhang, Yifeng Ding, Yongqiang Tian, Jianmin Guo 等ISSTA 2021 · 被引用 57 次
- Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language ModelsJingwei Yi, Yueqi Xie, Bin Zhu, Emre Kiciman 等KDD 2025 · 被引用 27 次
- MAWSEO: Adversarial Wiki Search Poisoning for Illicit Online PromotionZilong Lin, Zhengyi Li, Xiaojing Liao, XiaoFeng Wang 等S&P 2024 · 被引用 16 次
相关 Paper
- PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataTingchen Fu, Mrinank Sharma, Philip Torr, Shay B. Cohen 等ICML 2025
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang 等NDSS 2026 · 被引用 6 次
- Are LLM-Enhanced Graph Neural Networks Robust Against Poisoning Attacks?Yuhang Ma, Jie Wang, Zheng YanS&P 2026 · 被引用 4 次
- Prompt-to-SQL Injections in LLM-Integrated Web Applications: Risks and DefensesRodrigo Pedro, Miguel E. Coimbra, Daniel Castro, Paulo Carreira 等ICSE 2025 · 被引用 14 次
- PLeak: Prompt Leaking Attacks against Large Language Model ApplicationsBo Hui, Haolin Yuan, Neil Gong, Philippe Burlina 等CCS 2024 · 被引用 28 次
