AttnTrace: Contextual Attribution of Prompt Injection and Knowledge Corruption
Yanting Wang, Runpeng Geng, Ying Chen, Jinyuan Jia
Abstract
Long-context large language models (LLMs), such as GPT-5, Gemini-2.5-Pro, and Claude-Sonnet-4, are increasingly used to empower advanced AI systems, including retrievalaugmented generation (RAG) systems and autonomous agents. In these systems, an LLM receives an instruction along with a context-often consisting of texts retrieved from a knowledge database, memory, or the Internet-and generates a response that is contextually grounded by following the instruction. Many recent studies showed that these LLM-empowered systems are vulnerable to prompt injection and knowledge corruption attacks, where an attacker can inject malicious texts into the context such that the LLM generates an output as the attacker desires. One important research question is how to trace back to these malicious texts from a long context in leading to the attacker-desired output of the LLM. While significant efforts have been made, state-of-the-art solutions still achieve a sub-optimal performance and/or incur high computation cost. In this work, we propose AttnTrace, a new context traceback method based on the attention weights produced by an LLM for a prompt. To effectively utilize attention weights, we introduce two techniques designed to enhance the effectiveness of AttnTrace, and we provide theoretical insights for our design choice. We also perform a systematic evaluation for AttnTrace. The results demonstrate that AttnTrace is more accurate and efficient than existing state-of-the-art context traceback methods. We also show AttnTrace can improve state-of-the-art methods in detecting prompt injection under long contexts through the attribution-before-detection paradigm. As a real-world application, we demonstrate that AttnTrace can effectively pinpoint injected instructions in a paper designed to manipulate LLM-generated reviews. The code and data are available: https://github.com/Wang-Yanting/AttnTrace.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82432686-4164-4b0f-af91-eafc289fb7fbBuilds on32
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- TracLLM: A Generic Framework for Attributing Long Context LLMsYanting Wang, Wei Zou, Runpeng Geng, Jinyuan JiaUSENIX Security 2025
- Traceback of Poisoning Attacks to Retrieval-Augmented GenerationBaolei Zhang, Haoran Xin, Minghong Fang, Zhuqing Liu et al.WWW 2025 · 20 citations
- PR-Attack: Coordinated Prompt-RAG Attacks on Retrieval-Augmented Generation in Large Language Models via Bilevel OptimizationYang Jiao, Xiaodong Wang, Kai YangSIGIR 2025 · 6 citations
- Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation SystemsZhenting Qi, Hanlin Zhang, Eric P. Xing, Sham M. Kakade et al.ICLR 2025
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-SearchZeyu Shen, Basileal Imana, Tong Wu, Chong Xiang et al.NeurIPS 2025 · 26 citations
