Decoupling Knowledge and Context: An Efficient and Effective Retrieval Augmented Generation Framework via Cross Attention
Qian Dong, Qingyao Ai, Hongning Wang, Yiding Liu, Haitao Li, Weihang Su, Yiqun Liu, Tat-Seng Chua, Shaoping Ma
摘要
Retrieval-Augmented Generation (RAG) systems have become a crucial tool to augment large language models (LLMs) with external knowledge for better task performance. However, existing traditional RAG methods inject knowledge directly into the context, resulting in several limitations. First, these methods highly rely on the in-context learning capability of LLMs, which often leads to excessively long contexts. This is inefficient due to the quadratic complexity of self-attention, leading to significant increase in inference time. Second, the extended context and the nature of self-attention can cause the LLMs to lose important information in the context, thereby degrading the original capabilities of LLMs. Third, the effectiveness of knowledge injection is perturbed by the permutation of knowledge within the extended context, reducing the robustness of existing RAG methods. To tackle the above problems, we propose DecoupledRAG, a method that decouples external knowledge from the context within the RAG framework. Specifically, we introduce a cross-attention based method that injects retrieved knowledge directly into the inference process of LLM on the fly, without modifying its parameters or the input context, so that the external knowledge can be utilized robustly in a permutation-independent manner. To the best of our knowledge, this is the first work that explore how to utilize cross-attention to inject knowledge with low training cost in decoder-only LLM era. By leveraging cross-attention operation, DecoupledRAG enables seamless knowledge aggregation without creating extended context. Experimental results demonstrate that our method could achieve high efficiency while maintaining strong performance, which indicates that RAG frameworks have the potential to benefit further from more knowledge.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- Parametric Retrieval Augmented GenerationWeihang Su, Yichen Tang, Qingyao Ai, Junxi Yan 等SIGIR 2025 · 被引用 25 次
- Joint Evaluation of Answer and Reasoning Consistency for Hallucination Detection in Large Reasoning ModelsChangyue Wang, Weihang Su, Qingyao Ai, Yiqun LiuAAAI 2026 · 被引用 13 次
- Robust Fine-tuning for Retrieval Augmented Generation against Retrieval DefectsYiteng Tu, Weihang Su, Yujia Zhou, Yiqun Liu 等SIGIR 2025 · 被引用 9 次
- OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAGFengran Mo, Zhan Su, Yuchen Hui, Jinghan Zhang 等WWW 2026 · 被引用 8 次
- Bridging External and Parametric Knowledge: Mitigating Hallucination of LLMs with Shared-Private Semantic Synergy in Dual-Stream KnowledgeYi Sui, Chaozhuo Li, Chen Zhang, Dawei Song 等EMNLP 2025 · 被引用 1 次
相关 Paper
- UniRAG: Unified Query Understanding Method for Retrieval Augmented GenerationRui Li, Liyang He, Qi Liu, Zheng Zhang 等ACL 2025
- Incentivizing Retrieval-Augmented Generation via Inner Adaptive Context SelectionChenxu Cui, Lin Shen, Haihui Fan, Sa Zhu 等SIGIR 2026
- DRAGIN: Dynamic Retrieval Augmented Generation based on the Real-time Information Needs of Large Language ModelsWeihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu 等ACL 2024
- MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval AugmentationHongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao 等WWW 2025 · 被引用 92 次
- RAG+: Enhancing Retrieval-Augmented Generation with Application-Aware ReasoningYu Wang, Shiwan Zhao, Zhihu Wang, Ming Fan 等EMNLP 2025 · 被引用 3 次
