A Parametric Memory Head for Continual Generative Retrieval
Kidist Amde Mekonnen, Yubao Tang, Maarten de Rijke
Abstract
Generative information retrieval (GenIR) consolidates retrieval into a single neural model that decodes document identifiers (docids) directly from queries. While this model-as-index paradigm offers architectural simplicity, it is poorly suited to dynamic document collections. Unlike modular systems, where indexes are easily updated, GenIR's knowledge is parametrically encoded in its weights; consequently, standard adaptation methods such as full and parameter-efficient fine-tuning can induce catastrophic forgetting. We show that sequential adaptation improves retrieval on newly added documents but substantially degrades performance on earlier slices, exposing a pronounced stability-plasticity trade-off. To address this, we propose post-adaptation memory tuning (PAMT), a memory-only stabilization stage that augments an adapted model with a modular parametric memory head (PMH). PAMT freezes the backbone and attaches a product-key memory with fixed addressing. During prefix-trie constrained decoding, decoder hidden states sparsely query PMH to produce residual corrections in hidden space; these corrections are mapped to score adjustments via the frozen output embedding matrix, computed only over trie-valid tokens. This guides docid generation while keeping routing and backbone parameters fixed. To limit cross-slice interference, PAMT updates only a fixed budget of memory values selected using decoding-time access statistics, prioritizing entries frequently activated by the current slice and rarely used in prior sessions. Experiments on MS MARCO and Natural Questions under sequential, disjoint corpus increments show that PAMT substantially improves retention on earlier slices with minimal impact on retrieval performance for newly added documents, while modifying only a sparse subset of memory values per session.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6188261-cd00-44b8-b727-5e5141ec67a3Builds on18
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni et al.NeurIPS 2022 · 506 citations
- A Neural Corpus Indexer for Document RetrievalYujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao et al.NeurIPS 2022 · 242 citations
- Autoregressive Entity RetrievalNicola De Cao, Gautier Izacard, Sebastian Riedel, Fabio PetroniICLR 2021 · 200 citations
- Dense Passage Retrieval for Open-Domain Question AnsweringVladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis et al.EMNLP 2020 · 142 citations
Related papers
- Model Editing for New Document Integration in Generative Information RetrievalZhen Zhang, Zihan Wang, Xinyu Ma, Shuaiqiang Wang et al.WWW 2026
- DSI++: Updating Transformer Memory with New DocumentsSanket Vaibhav Mehta, Jai Gupta, Yi Tay, Mostafa Dehghani et al.EMNLP 2023 · 20 citations
- RAG without Forgetting: Continual Query-Infused Key MemoryYuntong Hu, Sha Li, Naren Ramakrishnan, Liang ZhaoICML 2026
- Lightweight and Direct Document Relevance Optimization for Generative Information RetrievalKidist Amde Mekonnen, Yubao Tang, Maarten de RijkeSIGIR 2025 · 3 citations
- TOME: A Two-stage Approach for Model-based RetrievalRuiyang Ren, Wayne Xin Zhao, Jing Liu, Hua Wu et al.ACL 2023 · 14 citations
