CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks
Xiaoxi Li, Zhicheng Dou, Yujia Zhou, Fangchao Liu
Abstract
Large language models (LLMs) have gained significant attention in various fields but prone to hallucination, especially in knowledge-intensive (KI) tasks. To address this, retrieval-augmented generation (RAG) has emerged as a popular solution to enhance factual accuracy. However, traditional retrieval modules often rely on large document index and disconnect with generative tasks. With the advent of generative retrieval (GR), language models can retrieve by directly generating document identifiers (DocIDs), offering superior performance in retrieval tasks. However, the potential relationship between GR and downstream tasks remains unexplored. In this paper, we propose CorpusLM, a unified language model that leverages external corpus to tackle various knowledge-intensive tasks by integrating generative retrieval, closed-book generation, and RAG through a unified greedy decoding process. We design the following mechanisms to facilitate effective retrieval and generation, and improve the end-to-end effectiveness of KI tasks: (1) We develop a ranking-oriented DocID list generation strategy, which refines GR by directly learning from a DocID ranking list, to improve retrieval quality. (2) We design a continuous DocIDs-References-Answer generation strategy, which facilitates effective and efficient RAG. (3) We employ well-designed unsupervised DocID understanding tasks, to comprehend DocID semantics and their relevance to downstream tasks. We evaluate our approach on the widely used KILT benchmark with two variants of backbone models, i.e., T5 and Llama2. Experimental results demonstrate the superior performance of our models in both retrieval and downstream tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1c8aab5-3820-4d8d-a0b7-5d8bf0ec328aCited by top-tier papers11
- HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG SystemsJiejun Tan, Zhicheng Dou, Wen Wang, Mang Wang et al.WWW 2025 · 42 citations
- Iterative Self-Incentivization Empowers Large Language Models as Agentic SearchersZhengliang Shi, Lingyong Yan, Dawei Yin, Suzan Verberne et al.NeurIPS 2025 · 15 citations
- Search-o1: Agentic Search-Enhanced Large Reasoning ModelsXiaoxi Li, Guanting Dong, Jiajie Jin, Yuyao Zhang et al.EMNLP 2025 · 12 citations
- BIRDIE: Natural Language-Driven Table Discovery Using Differentiable Search IndexYuxiang Guo, Zhonghao Hu, Yuren Mao, Baihua Zheng et al.VLDB 2025 · 6 citations
- On Synthetic Data Strategies for Domain-Specific Generative RetrievalHaoyang Wen, Jiang Guo, Yi Zhang, Jiarong Jiang et al.ACL 2025 · 6 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
Related papers
- Boosting Retrieval-Augmented Generation with Generation-Augmented Retrieval: A Co-Training ApproachYubao Tang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2025 · 2 citations
- RetroLLM: Empowering Large Language Models to Retrieve Fine-grained Evidence within GenerationXiaoxi Li, Jiajie Jin, Yujia Zhou, Yongkang Wu et al.ACL 2025
- Recitation-Augmented Language ModelsZhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang et al.ICLR 2023 · 30 citations
- A Unified Generative Retriever for Knowledge-Intensive Language Tasks via Prompt LearningJiangui Chen, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke et al.SIGIR 2023 · 31 citations
- UniRAG: Unified Query Understanding Method for Retrieval Augmented GenerationRui Li, Liyang He, Qi Liu, Zheng Zhang et al.ACL 2025
