Pre-computed memory or on-the-fly encoding? A hybrid approach to retrieval augmentation makes the most of your compute
Michiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald, Joshua Ainslie, Sumit Sanghai, Fei Sha, William W. Cohen
Abstract
Retrieval-augmented language models such as Fusion-in-Decoder are powerful, setting the state of the art on a variety of knowledge-intensive tasks. However, they are also expensive, due to the need to encode a large number of retrieved passages. Some work avoids this cost by pre-encoding a text corpus into a memory and retrieving dense representations directly. However, pre-encoding memory incurs a severe quality penalty as the memory representations are not conditioned on the current input. We propose LUMEN, a hybrid between these two extremes, pre-computing the majority of the retrieval representation and completing the encoding on the fly using a live encoder that is conditioned on the question and fine-tuned for the task. We show that LUMEN significantly outperforms pure memory on multiple question-answering tasks while being much cheaper than FiD, and outperforms both for any given compute budget. Moreover, the advantage of LUMEN over FiD increases with model size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5187fb68-30a7-4ef6-9403-3d47b5e1f161Cited by top-tier papers3
- Bridging the Preference Gap between Retrievers and LLMsZixuan Ke, Weize Kong, Cheng Li, Mingyang Zhang et al.ACL 2024 · 8 citations
- BTR: Binary Token Representations for Efficient Retrieval Augmented Language ModelsQingqing Cao, Sewon Min, Yizhong Wang, Hannaneh HajishirziICLR 2024 · 7 citations
- LAIT: Efficient Multi-Segment Encoding in Transformers with Layer-Adjustable InteractionJeremiah Milbauer, Annie Louis, Mohammad Javad Hosseini, Alex Fabrikant et al.ACL 2023 · 2 citations
Builds on16
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai et al.ICML 2022 · 1,629 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer et al.ICLR 2020 · 1,038 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
Related papers
- Optimizing Retrieval-augmented Reader Models via Token EliminationMoshe Berchansky, Peter Izsak, Avi Caciularu, Ido Dagan et al.EMNLP 2023 · 4 citations
- Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language ModelsJiaqi Cao, Jiarui Wang, Rubin Wei, Qipeng Guo et al.NeurIPS 2025 · 14 citations
- FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context LearningQinyuan Ye, Iz Beltagy, Matthew E. Peters, Xiang Ren et al.ACL 2023 · 5 citations
- Reusing Pre-Training Data at Test Time is a Compute MultiplierAlex Fang, Thomas Voice, Ruoming Pang, Ludwig Schmidt et al.ICLR 2026 · 4 citations
- MLP Memory: A Retriever-Pretrained Memory for Large Language ModelsRubin Wei, Jiaqi Cao, Jiarui Wang, Jushi Kai et al.ICLR 2026 · 16 citations
