Mnemosyne: Accelerating Multi-Hop Question Answering via Cache Hit Order Fitting
Haizhou Du, Jiujiu Li, Dongyang Li, Luobin Huang, Lisheng Wang
Abstract
Multi-Hop Question Answering (MHQA) requires stepby-step reasoning across multiple pieces of information to answer complex questions. The cache-aided Retrieval-Augmented Generation (RAG) can accelerate the process of external knowledge retrieval at each reasoning step for MHQA. However, existing methods focus on the internal structure and ignore the misalignment between the queries' arrival order and cache hit order. To tackle this, we propose Mnemosyne, a cache hit order fitting method designed to accelerate the RAG progress for MHQA. Specifically, our cache-aware order fitting strategy adjusts the order of queries arrival via graph reordering to better align with the cache hit order, thereby reducing the likelihood of failed or unproductive retrieval attempts. The multi-granularity caching storage mechanism is designed to loosen the strict hit condition to multiple similar semantic matching modes, facilitating that relevant documents can still be retrieved. Experiments conducted on four multi-hop QA datasets demonstrate that Mnemosyne effectively reduces retrieval latency while enhancing task answer F1 score, achieving a superior tradeoff between efficiency and effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60c41580-6146-4f74-8e6f-50019bfa1becBuilds on19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step QuestionsHarsh Trivedi, Niranjan Balasubramanian, Tushar Khot, Ashish SabharwalACL 2023 · 187 citations
- Extract, Define, Canonicalize: An LLM-based Framework for Knowledge Graph ConstructionBowen Zhang, Harold SohEMNLP 2024 · 65 citations
- Graph Reordering for Cache-Efficient Near Neighbor SearchBenjamin Coleman, Santiago Segarra, Alexander J. Smola, Anshumali ShrivastavaNeurIPS 2022 · 24 citations
Related papers
- MGRAG: Semantic Subgraph Matching and Graph-Aware Caching for Multimodal Retrieval-Augmented GenerationYubo Wang, Haoyang Li, Lei ChenVLDB 2026
- Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question AnsweringChangjian Wang, Weihong Deng, Weili Guan, Quan Lu et al.AAAI 2026 · 2 citations
- DualRAG: A Dual-Process Approach to Integrate Reasoning and Retrieval for Multi-Hop Question AnsweringRong Cheng, Jinyi Liu, Yan Zheng, Fei Ni et al.ACL 2025
- Question-Adaptive Graph Learning for Multi-hop Retrieval Augmented GenerationYuchen Yan, Peiyan Zhang, Zhihua Liu, Hao Wang et al.SIGIR 2026
- Iterative Multi-Granular RAG with Contextual Hierarchical GraphYanli Hu, Teng Liu, Zhuangyi Zhou, Weixin Zeng et al.AAAI 2026
