SAGE: A Framework of Precise Retrieval for RAG
Jintao Zhang, Guoliang Li, Jinyang Su
Abstract
Retrieval-augmented generation (RAG) has demonstrated significant proficiency in conducting question-answering (QA) tasks within a specified corpus. Nonetheless, numerous failure instances of RAG in QA still exist. These failures are not solely attributable to the limitations of Large Language Models (LLMs); instead, they predominantly arise from the retrieval of inaccurate information for LLMs due to two limitations: (1) Current RAG methods segment the corpus without considering semantics, making it difficult to find relevant context due to impaired correlation between questions and the segments. (2) There's a trade-off between missing essential context with fewer context retrieved and getting irrelevant context with more context retrieved. It is hard to make an ideal balance.
In this paper, we introduce a RAG framework, named SAGE, designed to overcome these limitations. First, to address the issue of segmentation without considering semantics, we propose to train a semantic segmentation model. This model is trained to segment the corpus into semantically complete chunks. Second, to ensure that only the most relevant chunks are retrieved while the irrelevant ones are ignored, we design a chunk selection algorithm to dynamically select chunks based on the decreasing speed of the relevance score of chunks, leading to a more relevant selection. Third, to further ensure the precision of the retrieved chunks, we propose letting LLMs assess whether retrieved chunks are excessive or lacking and then adjust the amount of context accordingly. Experimental results show that SAGE outperforms baselines by 61.25% in the quality of QA on average. Moreover, by avoiding retrieving noisy context, SAGE lowers the cost of the tokens consumed in LLM inference and achieves a 49.41% enhancement in cost efficiency on average. Additionally, our work offers valuable insights for boosting RAG, contributing to the development of more effective RAG systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5b0682f9-c002-4bf8-82c7-55fb50c31eb6Cited by top-tier papers4
- XRAG: Examining the Core - Benchmarking Foundational Components in Advanced Retrieval-Augmented GenerationQili Zhang, Qianren Mao, Yangyifei Luo, Yashuo Luo et al.ICDE 2026 · 1 citation
- SageAttention: Accurate 8-Bit Attention for Plug-and-play Inference AccelerationJintao Zhang, Jia Wei, Pengle Zhang, Jun Zhu et al.ICLR 2025
- SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model InferenceJintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei et al.ICML 2025
- SageAttention2: Efficient Attention with Thorough Outlier Smoothing and Per-thread INT4 QuantizationJintao Zhang, Haofeng Huang, Pengle Zhang, Jia Wei et al.ICML 2025
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna et al.ICLR 2024 · 460 citations
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- The Power of Noise: Redefining Retrieval for RAG SystemsFlorin Cuconasu, Giovanni Trappolini, Federico Siciliano, Simone Filice et al.SIGIR 2024 · 212 citations
Related papers
- Accelerating Inference of Retrieval-Augmented Generation via Sparse Context SelectionYun Zhu, Jia-Chen Gu, Caitlin Sikora, Ho Ko et al.ICLR 2025
- RetroLM: Retrieval-Augmented KVs for Long-Context ProcessingKun Luo, Zheng Liu, Shitao Xiao, Jiabei Chen et al.AAAI 2026
- From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented GenerationJiahao Wang, Weiyu Xie, Mingxing Zhang, Boxin Zhang et al.SIGMOD 2026 · 4 citations
- AdaCache: Adaptive Caching and Context Augmentation for Efficient LLM ServingZihao Zeng, Siyi Li, Xinyu Yan, Lei Xiao et al.ICLR 2026
- You Don't Need Pre-Built Graphs for RAG: Retrieval Augmented Generation with Adaptive Reasoning StructuresShengyuan Chen, Chuang Zhou, Zheng Yuan, Qinggang Zhang et al.AAAI 2026 · 14 citations
