SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
Nan Zhang, Prafulla Kumar Choubey, Alexander R. Fabbri, Gabriel Bernadett-Shapiro, Rui Zhang, Prasenjit Mitra, Caiming Xiong, Chien-Sheng Wu
Abstract
Indexing is an important step towards strong performance in retrieval-augmented generation (RAG) systems. However, existing methods organize data based on either semantic similarity (similarity) or related information (relatedness), but do not cover both perspectives comprehensively. Our analysis reveals that modeling only one perspective results in insufficient knowledge synthesis, leading to suboptimal performance on complex tasks requiring multihop reasoning. In this paper, we propose SiReRAG, a novel RAG indexing approach that explicitly considers both similar and related information. On the similarity side, we follow existing work and explore some variances to construct a similarity tree based on recursive summarization. On the relatedness side, SiReRAG extracts propositions and entities from texts, groups propositions via shared entities, and generates recursive summaries to construct a relatedness tree. We index and flatten both similarity and relatedness trees into a unified retrieval pool. Our experiments demonstrate that SiReRAG consistently outperforms state-of-the-art indexing methods on three multihop datasets (MuSiQue, 2WikiMultiHopQA, and HotpotQA), with an average 1.9% improvement in F1 scores. As a reasonably efficient solution, SiReRAG enhances existing reranking methods significantly, with up to 7.8% improvement in average F1 scores. Our code is available at https://github.com/SalesforceAIResearch/SiReRAG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f39a1d5a-a5dd-4ede-a912-ce883bba46aaCited by top-tier papers10
- G-reasoner: Foundation Models for Unified Reasoning over Graph-structured KnowledgeLinhao Luo, Zicheng Zhao, Junnan Liu, Zhangchi Qiu et al.ICLR 2026 · 12 citations
- SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language ModelsKen Gu, Advait Bhat, Mike A Merrill, Robert West et al.ICLR 2026 · 6 citations
- When Reasoning Meets Compression: Understanding the Effects of LLMs Compression on Large Reasoning ModelsNan Zhang, Eugene Kwek, Yusen Zhang, Hieu Nguyen et al.ICLR 2026 · 5 citations
- AtlasKV: Augmenting LLMs with Billion-Scale Knowledge Graphs in 20GB VRAMHaoyu Huang, Hong Ting Tsang, Jiaxin Bai, Xi Peng et al.ICLR 2026 · 4 citations
- HyperMem: Hypergraph Memory for Long-Term ConversationsJuwei Yue, Chuanrui Hu, Jiawei Sheng, Zuyi Zhou et al.ACL 2026 · 4 citations
Builds on9
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna et al.ICLR 2024 · 460 citations
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga et al.NeurIPS 2024 · 395 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
- Generate rather than Retrieve: Large Language Models are Strong Context GeneratorsWenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu et al.ICLR 2023 · 86 citations
Related papers
- NeocorRAG: Less Irrelevant Information, More Explicit Evidence, and More Effective Recall via Evidence ChainsShiyao Peng, Qianhe Zheng, Zhuodi Hao, Zichen Tang et al.WWW 2026
- PropRAG: Guiding Retrieval with Beam Search over Proposition PathsJingjin Wang, Jiawei HanEMNLP 2025
- CIRAG: Retrieval-Augmented Language Model with Collective IntelligenceChenxu Cui, Haihui Fan, Jinchao Zhang, Lin Shen et al.SIGIR 2025 · 5 citations
- HyperRAG: Reasoning N-ary Facts over Hypergraphs for Retrieval Augmented GenerationWen-Sheng Lien, Yu-Kai Chan, Hao-Lung Hsiao, Bo-Kai Ruan et al.WWW 2026
- Token-Free Hierarchical Indexing for RAG beyond LLM-based SummarizationYifan Wei, Dan Yuan, Xiaoyan Yu, Angsheng LiICML 2026
