Generative Retrieval via Term Set Generation
Peitian Zhang, Zheng Liu, Yujia Zhou, Zhicheng Dou, Fangchao Liu, Zhao Cao
Abstract
Recently, generative retrieval has emerged as a promising alternative to the traditional retrieval paradigms. It assigns each document a unique identifier, known as the DocID, and employs a generative model to directly generate the relevant DocID for the input query. A common choice for the DocID is one or several natural language sequences, e.g. the title, synthetic queries, or n-grams, so that the pre-trained knowledge of the generative model can be effectively utilized. However, a sequence is generated token by token, where only the most likely candidates are kept and the rest are pruned at each decoding step, thus, retrieval fails if any token within the relevant DocID is falsely pruned. What's worse, during decoding, the model can only perceive preceding tokens in the DocID while being blind to subsequent ones, hence is prone to make such errors. To address this problem, we present a novel framework for generative retrieval, dubbed Term-Set Generation (TSGen). Instead of sequences, we use a set of terms as the DocID. The terms are selected based on learned weights from relevance signals, so that they concisely summarize the document's semantics and distinguish it from others. On top of the term-set DocID, we propose a permutation-invariant decoding algorithm, with which the term set can be generated in any permutation yet will always lead to the corresponding document. Remarkably, TSGen perceives all valid terms rather than only the preceding ones at each decoding step. Given the constant decoding space, it can make more reliable decisions due to the broader perspective. TSGen is also resilient to errors: the relevant DocID will not be falsely pruned as long as the decoded term belongs to it. Moreover, TSGen can explore the optimal decoding permutation of the term set on its own, which further improves the likelihood of generating the relevant DocID. Lastly, we design an iterative optimization procedure to incentivize the model to generate the relevant term set in its favorable permutation. We conduct extensive experiments on popular benchmarks of generative retrieval, which validate the effectiveness, the generalizability, the scalability, and the efficiency of TSGen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Generative Retrieval Meets Multi-Graded RelevanceYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.NeurIPS 2024 · 18 citations
- CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive TasksXiaoxi Li, Zhicheng Dou, Yujia Zhou, Fangchao LiuSIGIR 2024 · 16 citations
- Order-agnostic Identifier for Large Language Model-based Generative RecommendationXinyu Lin, Haihan Shi, Wenjie Wang, Fuli Feng et al.SIGIR 2025 · 15 citations
- On Synthetic Data Strategies for Domain-Specific Generative RetrievalHaoyang Wen, Jiang Guo, Yi Zhang, Jiarong Jiang et al.ACL 2025 · 6 citations
- Descriptive and Discriminative Document Identifiers for Generative RetrievalJiehan Cheng, Zhicheng Dou, Yutao Zhu, Xiaoxi LiAAAI 2025 · 4 citations
Builds on24
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang et al.ICLR 2021 · 1,547 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
- Recommender Systems with Generative RetrievalShashank Rajput, Nikhil Mehta, Anima Singh, Raghunandan Hulikal Keshavan et al.NeurIPS 2023 · 474 citations
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
- A Neural Corpus Indexer for Document RetrievalYujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao et al.NeurIPS 2022 · 242 citations
Related papers
- GLEN: Generative Retrieval via Lexical Index LearningSunkyung Lee, Minjin Choi, Jongwuk LeeEMNLP 2023 · 6 citations
- Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous DecodingHansi Zeng, Chen Luo, Hamed ZamaniSIGIR 2024 · 21 citations
- Multi-level Relevance Document Identifier Learning for Generative RetrievalFuwei Zhang, Xiaoyu Liu, Xinyu Jia, Yingfei Zhang et al.ACL 2025 · 5 citations
- Enhancing Generative Retrieval with Reinforcement Learning from Relevance FeedbackYujia Zhou, Zhicheng Dou, Ji-Rong WenEMNLP 2023 · 14 citations
- Learning to Tokenize for Generative RetrievalWeiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang et al.NeurIPS 2023 · 151 citations
