DSI++: Updating Transformer Memory with New Documents
Sanket Vaibhav Mehta, Jai Gupta, Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Jinfeng Rao, Marc Najork, Emma Strubell, Donald Metzler
摘要
Differentiable Search Indices (DSIs) encode a corpus of documents in model parameters and use the same model to answer user queries directly. Despite the strong performance of DSI models, deploying them in situations where the corpus changes over time is computationally expensive because reindexing the corpus requires re-training the model. In this work, we introduce DSI++, a continual learning challenge for DSI to incrementally index new documents while being able to answer queries related to both previously and newly indexed documents. Across different model scales and document identifier representations, we show that continual indexing of new documents leads to considerable forgetting of previously indexed documents. We also hypothesize and verify that the model experiences forgetting events during training, leading to unstable learning. To mitigate these issues, we investigate two approaches. The first focuses on modifying the training dynamics. Flatter minima implicitly alleviate forgetting, so we optimize for flatter loss basins and show that the model stably memorizes more documents (+12%). Next, we introduce a generative memory to sample pseudoqueries for documents and supplement them during continual indexing to prevent forgetting for the retrieval task. Extensive experiments on novel continual indexing benchmarks based on Natural Questions (NQ) and MS MARCO demonstrate that our proposed solution mitigates forgetting significantly. Concretely, it improves the average Hits@10 by +21.1% over competitive baselines for NQ and requires 6 times fewer model updates compared to retraining the DSI model for incrementally indexing five corpora in a sequence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Learning to Tokenize for Generative RetrievalWeiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang 等NeurIPS 2023 · 被引用 151 次
- Scalable and Effective Generative Information RetrievalHansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar 等WWW 2024 · 被引用 72 次
- UniGen: A Unified Generative Framework for Retrieval and Question Answering with Large Language ModelsXiaoxi Li, Yujia Zhou, Zhicheng DouAAAI 2024 · 被引用 24 次
- How Does Generative Retrieval Scale to Millions of Passages?Ronak Pradeep, Kai Hui, Jai Gupta, Ádám D. Lelkes 等EMNLP 2023 · 被引用 23 次
- Planning Ahead in Generative Retrieval: Guiding Autoregressive Generation through Simultaneous DecodingHansi Zeng, Chen Luo, Hamed ZamaniSIGIR 2024 · 被引用 21 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Fast Model Editing at ScaleEric Mitchell, Charles Lin, Antoine Bosselut, Chelsea Finn 等ICLR 2022 · 被引用 527 次
相关 Paper
- IncDSI: Incrementally Updatable Document RetrievalVarsha Kishore, Chao Wan, Justin Lovelace, Yoav Artzi 等ICML 2023 · 被引用 19 次
- Transformer Memory as a Differentiable Search IndexYi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni 等NeurIPS 2022 · 被引用 506 次
- A Parametric Memory Head for Continual Generative RetrievalKidist Amde Mekonnen, Yubao Tang, Maarten de RijkeSIGIR 2026
- Exploring the Practicality of Generative Retrieval on Dynamic CorporaChaeeun Kim, Soyoung Yoon, Hyunji Lee, Joel Jang 等EMNLP 2024 · 被引用 1 次
- Panini: Continual Learning in Token Space via Structured MemoryShreyas Rajesh, Pavan Holur, Mehmet Yigit Turali, Chenda Duan 等ICML 2026 · 被引用 1 次
