How Does Generative Retrieval Scale to Millions of Passages?
Ronak Pradeep, Kai Hui, Jai Gupta, Ádám D. Lelkes, Honglei Zhuang, Jimmy Lin, Donald Metzler, Vinh Q. Tran
Abstract
The emerging paradigm of generative retrieval re-frames the classic information retrieval problem into a sequence-to-sequence modeling task, forgoing external indices and encoding an entire document corpus within a single Transformer. Although many different approaches have been proposed to improve the effectiveness of generative retrieval, they have only been evaluated on document corpora on the order of 100K in size. We conduct the first empirical study of generative retrieval techniques across various corpus scales, ultimately scaling up to the entire MS MARCO passage ranking task with a corpus of 8.8M passages and evaluating model sizes up to 11B parameters. We uncover several findings about scaling generative retrieval to millions of passages; notably, the central importance of using synthetic queries as document representations during indexing, the ineffectiveness of existing proposed architecture modifications when accounting for compute cost, and the limits of naively scaling model parameters with respect to retrieval performance. While we find that generative retrieval is competitive with stateof-the-art dual encoders on small corpora, scaling to millions of passages remains an important and unsolved challenge. We believe these findings will be valuable for the community to clarify the current state of generative retrieval, highlight the unique challenges, and inspire new research directions. * Equal Contribution. † Work completed while a Student Researcher at Google.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7e0b8ff8-1601-422e-9199-7a04b311cb48Cited by top-tier papers20
- Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous DecodingHansi Zeng, Chen Luo, Hamed ZamaniSIGIR 2024 · 21 citations
- Generative Retrieval Meets Multi-Graded RelevanceYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.NeurIPS 2024 · 18 citations
- Prompt Perturbation in Retrieval-Augmented Generation based Large Language ModelsZhibo Hu, Chen Wang, Yanfeng Shu, Hye-Young Paik et al.KDD 2024 · 15 citations
- Learning Facts at Scale with Active ReadingJessy Lin, Vincent-Pierre Berges, Xilun Chen, Wen-tau Yih et al.ICLR 2026 · 15 citations
- Self-Retrieval: End-to-End Information Retrieval with One Large Language ModelQiaoyu Tang, Jiawei Chen, Zhuoqun Li, Bowen Yu et al.NeurIPS 2024 · 14 citations
Builds on10
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 2,496 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
- A Neural Corpus Indexer for Document RetrievalYujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao et al.NeurIPS 2022 · 242 citations
Related papers
- Exploring Training and Inference Scaling Laws in Generative RetrievalHongru Cai, Yongqi Li, Ruifeng Yuan, Wenjie Wang et al.SIGIR 2025 · 1 citation
- Scalable and Effective Generative Information RetrievalHansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar et al.WWW 2024 · 72 citations
- GLEN: Generative Retrieval via Lexical Index LearningSunkyung Lee, Minjin Choi, Jongwuk LeeEMNLP 2023 · 6 citations
- Multiview Identifiers Enhanced Generative RetrievalYongqi Li, Nan Yang, Liang Wang, Furu Wei et al.ACL 2023 · 30 citations
- Multi-stage Training with Improved Negative Contrast for Neural Passage RetrievalJing Lu, Gustavo Hernández Ábrego, Ji Ma, Jianmo Ni et al.EMNLP 2021 · 21 citations
