AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
Yanting Wang, Runpeng Geng, Ying Chen, Jinyuan Jia
摘要
Long-context large language models (LLMs), such as GPT-5, Gemini-2.5-Pro, and Claude-Sonnet-4, are increasingly used to empower advanced AI systems, including retrievalaugmented generation (RAG) systems and autonomous agents. In these systems, an LLM receives an instruction along with a context-often consisting of texts retrieved from a knowledge database, memory, or the Internet-and generates a response that is contextually grounded by following the instruction. Many recent studies showed that these LLM-empowered systems are vulnerable to prompt injection and knowledge corruption attacks, where an attacker can inject malicious texts into the context such that the LLM generates an output as the attacker desires. One important research question is how to trace back to these malicious texts from a long context in leading to the attacker-desired output of the LLM. While significant efforts have been made, state-of-the-art solutions still achieve a sub-optimal performance and/or incur high computation cost. In this work, we propose AttnTrace, a new context traceback method based on the attention weights produced by an LLM for a prompt. To effectively utilize attention weights, we introduce two techniques designed to enhance the effectiveness of AttnTrace, and we provide theoretical insights for our design choice. We also perform a systematic evaluation for AttnTrace. The results demonstrate that AttnTrace is more accurate and efficient than existing state-of-the-art context traceback methods. We also show AttnTrace can improve state-of-the-art methods in detecting prompt injection under long contexts through the attribution-before-detection paradigm. As a real-world application, we demonstrate that AttnTrace can effectively pinpoint injected instructions in a paper designed to manipulate LLM-generated reviews. The code and data are available: https://github.com/Wang-Yanting/AttnTrace.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- TracLLM: A Generic Framework for Attributing Long Context LLMsYanting Wang, Wei Zou, Runpeng Geng, Jinyuan JiaUSENIX Security 2025
- Traceback of Poisoning Attacks to Retrieval-Augmented GenerationBaolei Zhang, Haoran Xin, Minghong Fang, Zhuqing Liu 等WWW 2025 · 被引用 20 次
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 被引用 6 次
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation SystemsZhenting Qi, Hanlin Zhang, Eric P. Xing, Sham M. Kakade 等ICLR 2025
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-SearchZeyu Shen, Basileal Imana, Tong Wu, Chong Xiang 等NeurIPS 2025 · 被引用 26 次
