Generative Multi-hop Retrieval
Hyunji Lee, Sohee Yang, Hanseok Oh, Minjoon Seo
Abstract
A common practice for text retrieval is to use an encoder to map the documents and the query to a common vector space and perform a nearest neighbor search (NNS); multi-hop retrieval also often adopts the same paradigm, usually with a modification of iteratively reformulating the query vector so that it can retrieve different documents at each hop. However, such a biencoder approach has limitations in multi-hop settings; (1) the reformulated query gets longer as the number of hops increases, which further tightens the embedding bottleneck of the query vector, and (2) it is prone to error propagation. In this paper, we focus on alleviating these limitations in multi-hop settings by formulating the problem in a fully generative way. We propose an encoder-decoder model that performs multi-hop retrieval by simply generating the entire text sequences of the retrieval targets, which means the query and the documents interact in the language model's parametric space rather than L2 or inner product space as in the bi-encoder approach. Our approach, Generative Multi-hop Retrieval (GMR), consistently achieves comparable or higher performance than bi-encoder models in five datasets while demonstrating superior GPU memory and storage footprint. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Precise Zero-Shot Dense Retrieval without Relevance LabelsLuyu Gao, Xueguang Ma, Jimmy Lin, Jamie CallanACL 2023 · 211 citations
- Synergistic Interplay between Search and Large Language Models for Information RetrievalJiazhan Feng, Chongyang Tao, Xiubo Geng, Tao Shen et al.ACL 2024 · 7 citations
- MRR-FV: Unlocking Complex Fact Verification with Multi-Hop Retrieval and ReasoningLiwen Zheng, Chaozhuo Li, Litian Zhang, Haoran Jia et al.AAAI 2025 · 6 citations
- On Synthetic Data Strategies for Domain-Specific Generative RetrievalHaoyang Wen, Jiang Guo, Yi Zhang, Jiarong Jiang et al.ACL 2025 · 6 citations
- Generative Subgraph Retrieval for Knowledge Graph-Grounded Dialog GenerationJinyoung Park, Minseok Joo, Joo-Kyung Kim, Hyunwoo J. KimEMNLP 2024 · 3 citations
Builds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Scalable Zero-shot Entity Linking with Dense Entity RetrievalLedell Wu, Fabio Petroni, Martin Josifoski, Sebastian Riedel et al.EMNLP 2020 · 336 citations
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
- Answering Complex Open-Domain Questions with Multi-Hop Dense RetrievalWenhan Xiong, Xiang Lorraine Li, Srini Iyer, Jingfei Du et al.ICLR 2021 · 232 citations
Related papers
- How Does Generative Retrieval Scale to Millions of Passages?Ronak Pradeep, Kai Hui, Jai Gupta, Ádám D. Lelkes et al.EMNLP 2023 · 23 citations
- Scaling the Vocabulary of Non-autoregressive Models for Fast Generative RetrievalRavisri Valluri, Akash Kumar Mohankumar, Kushal Dave, Amit Singh et al.KDD 2025
- Making Transformer Decoders Better Differentiable IndexersWuchao Li, Kai Zheng, Defu Lian, Qi Liu et al.ICLR 2025
- Large Dual Encoders Are Generalizable RetrieversJianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai et al.EMNLP 2022 · 145 citations
- LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query InferenceGuangyuan Ma, Yongliang Ma, Xuanrui Gou, Zhenpeng Su et al.ICLR 2026 · 3 citations
