xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
Xin Cheng, Xun Wang, Xingxing Zhang, Tao Ge, Si-Qing Chen, Furu Wei, Huishuai Zhang, Dongyan Zhao
Abstract
This paper introduces xRAG, a novel context compression method designed specifically for retrieval-augmented generation. xRAG redefines the use of document embeddings in dense retrieval-traditionally limited to retrieval purposes-by integrating them as features from the retrieval modality. Through a modality fusion approach, xRAG effectively merges these embeddings into the language model's representation space, eliminating the need for their textual counterparts and achieving an extreme compression rate. In xRAG, the modality bridge is the only trainable component, while the retriever and language model remain frozen. This design choice allows for the reuse of offline-constructed document embeddings and preserves the plug-and-play nature of retrieval augmentation. Experimental results demonstrate that xRAG achieves an average improvement of over 10% across six knowledge-intensive tasks, compatible with various language model backbones, ranging from a dense 7B model to an 8x7B Mixture of Experts configuration. xRAG not only significantly outperforms previous context compression methods but also matches the performance of uncompressed models on several benchmarks, while reducing overall FLOPs by a factor of 3.53. This work pioneers new avenues in retrieval-augmented generation through multimodal fusion, potentially setting a groundwork for future developments in efficient and scalable retrieval systems. How might we mitigate the costs associated with extended context while maintaining the benefits of retrieval augmentation? Recent research interest has converged on a promising direction: Context Compression. This concept is pursued through two primary strategies: soft-prompting methods, such as Gist [58], AutoCompressor [14], and ICAE [19] , which compress the context into dense memory slots, and hard-prompting methods, such as LLMLingua [28] and RECOMP [79] , where
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02cf8126-ccc0-4ac1-b983-077af59cd92fCited by top-tier papers30
- 500xCompressor: Generalized Prompt Compression for Large Language ModelsZongqian Li, Yixuan Su, Nigel CollierACL 2025 · 35 citations
- Leveraging Passage Embeddings for Efficient Listwise Reranking with Large Language ModelsQi Liu, Bo Wang, Nan Wang, Jiaxin MaoWWW 2025 · 26 citations
- PromptReps: Prompting Large Language Models to Generate Dense and Sparse Representations for Zero-Shot Document RetrievalShengyao Zhuang, Xueguang Ma, Bevan Koopman, Jimmy Lin et al.EMNLP 2024 · 26 citations
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye et al.ACL 2026 · 24 citations
- What Generative Search Engines Like and How to Optimize Web Content CooperativelyYujiang Wu, Shanshan Zhong, Yubin Kim, Chenyan XiongICLR 2026 · 19 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- CompAct: Compressing Retrieved Documents Actively for Question AnsweringChanwoong Yoon, Taewhoo Lee, Hyeon Hwang, Minbyul Jeong et al.EMNLP 2024 · 10 citations
- Less Is More: Elevating RAG via Performance-Driven Context CompressionZiqiang Cui, Yunpeng Weng, Xing Tang, Peiyang Liu et al.ICML 2026
- Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector PerspectiveYunhao Liu, Zian Jia, Xinyu Gao, Kanjun Xu et al.WWW 2026
- Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionYun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko et al.ICLR 2025
- FlowRAG: Continual Learning for Dynamic Retriever in Retrieval-Augmented GenerationSenlei Zhang, Tongjun Shi, Dandan Song, Luan Zhang et al.WWW 2026
