A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt Learning
Jiangui Chen, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yiqun Liu, Yixing Fan, Xueqi Cheng
Abstract
Knowledge-intensive language tasks (KILTs) benefit from retrieving high-quality relevant contexts from large external knowledge corpora. Learning task-specific retrievers that return relevant contexts at an appropriate level of semantic granularity, such as a document retriever, passage retriever, sentence retriever, and entity retriever, may help to achieve better performance on the end-to-end task. But a task-specific retriever usually has poor generalization ability to new domains and tasks, and it may be costly to deploy a variety of specialised retrievers in practice.
We propose a unified generative retriever (UGR) that combines task-specific effectiveness with robust performance over different retrieval tasks in KILTs. To achieve this goal, we make two major contributions: (i) To unify different retrieval tasks into a single generative form, we introduce an n-gram-based identifier for relevant contexts at different levels of granularity in KILTs. And (ii) to address different retrieval tasks with a single model, we employ a prompt learning strategy and investigate three methods to design prompt tokens for each task. In this way, the proposed UGR model can not only share common knowledge across tasks for better generalization, but also perform different retrieval tasks effectively by distinguishing task-specific characteristics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Generative Multi-Modal Knowledge Retrieval with Large Language ModelsXinwei Long, Jiali Zeng, Fandong Meng, Zhiyuan Ma et al.AAAI 2024 · 35 citations
- Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous DecodingHansi Zeng, Chen Luo, Hamed ZamaniSIGIR 2024 · 21 citations
- Generative Retrieval Meets Multi-Graded RelevanceYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.NeurIPS 2024 · 18 citations
- Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational SearchYifei Yuan, Clemencia Siro, Mohammad Aliannejadi, Maarten de Rijke et al.WWW 2024 · 17 citations
- CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive TasksXiaoxi Li, Zhicheng Dou, Yujia Zhou, Fangchao LiuSIGIR 2024 · 16 citations
Builds on13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
Related papers
- A Prompt-Based Knowledge Graph Foundation Model for Universal In-Context ReasoningYuanning Cui, Zequn Sun, Wei HuNeurIPS 2024 · 46 citations
- UniLP: Unified Topology-aware Generative Framework for Link Prediction in Knowledge GraphBen Liu, Miao Peng, Wenjie Xu, Xu Jia et al.WWW 2024 · 9 citations
- HypeR: Multitask Hyper-Prompted Training Enables Large-Scale Retrieval GeneralizationZefeng Cai, Chongyang Tao, Tao Shen, Can Xu et al.ICLR 2023
- Training a Utility-based Retriever Through Shared Context Attribution for Retrieval-Augmented Language ModelsYilong Xu, Jinhua Gao, Xiaoming Yu, Yuanhai Xue et al.EMNLP 2025
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih et al.NeurIPS 2022 · 242 citations
