UniGen: A Unified Generative Framework for Retrieval and Question Answering with Large Language Models
Xiaoxi Li, Yujia Zhou, Zhicheng Dou
摘要
Generative information retrieval, encompassing two major tasks of Generative Document Retrieval (GDR) and Grounded Answer Generation (GAR), has gained significant attention in the area of information retrieval and natural language processing. Existing methods for GDR and GAR rely on separate retrieval and reader modules, which hinder simultaneous optimization. To overcome this, we present UniGen, a Unified Generative framework for retrieval and question answering that integrates both tasks into a single generative model leveraging the capabilities of large language models. UniGen employs a shared encoder and two distinct decoders for generative retrieval and question answering. To facilitate the learning of both tasks, we introduce connectors, generated by large language models, to bridge the gaps between query inputs and generation targets, as well as between document identifiers and answers. Furthermore, we propose an iterative enhancement strategy that leverages generated answers and retrieved documents to iteratively improve both tasks. Through extensive experiments on the MS MARCO and NQ datasets, we demonstrate the effectiveness of UniGen, showcasing its superior performance in both the retrieval and the question answering tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Plan-on-Graph: Self-Correcting Adaptive Planning of Large Language Model on Knowledge GraphsLiyi Chen, Panrong Tong, Zhongming Jin, Ying Sun 等NeurIPS 2024 · 被引用 160 次
- HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsJiejun Tan, Zhicheng Dou, Wen Wang, Mang Wang 等WWW 2025 · 被引用 42 次
- Metacognitive Retrieval-Augmented Large Language ModelsYujia Zhou, Zheng Liu, Jiajie Jin, Jian-Yun Nie 等WWW 2024 · 被引用 41 次
- DeepAgent: A General Reasoning Agent with Scalable ToolsetsXiaoxi Li, Wenxiang Jiao, Jiarui Jin, Guanting Dong 等WWW 2026 · 被引用 38 次
- CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive TasksXiaoxi Li, Zhicheng Dou, Yujia Zhou, Fangchao LiuSIGIR 2024 · 被引用 16 次
它引用的顶会 Paper16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text RetrievalLee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang 等ICLR 2021 · 被引用 1,547 次
- Autoregressive Search Engines: Generating Substrings as Document IdentifiersMichele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih 等NeurIPS 2022 · 被引用 242 次
相关 Paper
- Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question AnsweringZhengliang Shi, Shuo Zhang, Weiwei Sun, Shen Gao 等ACL 2024
- Generate rather than Retrieve: Large Language Models are Strong Context GeneratorsWenhao Yu, Dan Iter, Shuohang Wang, Yichong Xu 等ICLR 2023 · 被引用 86 次
- UniRAG: Unified Query Understanding Method for Retrieval Augmented GenerationRui Li, Liyang He, Qi Liu, Zheng Zhang 等ACL 2025
- List-aware Reranking-Truncation Joint Model for Search and Retrieval-augmented GenerationShicheng Xu, Liang Pang, Jun Xu, Huawei Shen 等WWW 2024 · 被引用 13 次
- Harnessing Multi-Role Capabilities of Large Language Models for Open-Domain Question AnsweringHongda Sun, Yuxuan Liu, Chengwei Wu, Haiyu Yan 等WWW 2024 · 被引用 16 次
