Memorizing Transformers
Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins, Christian Szegedy
摘要
Language models typically need to be trained or finetuned in order to acquire new knowledge, which involves updating their weights. We instead envision language models that can simply read and memorize new data at inference time, thus acquiring new knowledge immediately. In this work, we extend language models with the ability to memorize the internal representations of past inputs. We demonstrate that an approximate kNN lookup into a non-differentiable memory of recent (key, value) pairs improves language modeling across various benchmarks and tasks, including generic webtext (C4), math papers (arXiv), books (PG-19), code (Github), as well as formal theorems (Isabelle). We show that the performance steadily improves when we increase the size of memory up to 262K tokens. On benchmarks including code and mathematics, we find that the model is capable of making use of newly defined functions and theorems during test time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper104
- Learning to Compress Prompts with Gist TokensJesse Mu, Xiang Li, Noah D. GoodmanNeurIPS 2023 · 被引用 488 次
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth ApproachJonas Geiping, Sean McLeish, Neel Jain, John Kirchenbauer 等NeurIPS 2025 · 被引用 431 次
- Generating Images with Multimodal Language ModelsJing Yu Koh, Daniel Fried, Russ SalakhutdinovNeurIPS 2023 · 被引用 403 次
- Confident Adaptive Language ModelingTal Schuster, Adam Fisch, Jai Gupta, Mostafa Dehghani 等NeurIPS 2022 · 被引用 394 次
- Autoformalization with Large Language ModelsYuhuai Wu, Albert Qiaochu Jiang, Wenda Li, Markus N. Rabe 等NeurIPS 2022 · 被引用 364 次
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
相关 Paper
- FictionalQA: A Dataset for Studying Memorization and Knowledge AcquisitionJohn Kirchenbauer, Natjanan Mongkolsupawan, Yuxin Wen, Tom Goldstein 等ICLR 2026 · 被引用 1 次
- Pre-training Limited Memory Language Models with Internal and External KnowledgeLinxi Zhao, Sofian Zalouk, Christian K. Belardi, Justin Lovelace 等ICLR 2026 · 被引用 11 次
- Why do Nearest Neighbor Language Models Work?Frank F. Xu, Uri Alon, Graham NeubigICML 2023 · 被引用 33 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- MEMORYLLM: Towards Self-Updatable Large Language ModelsYu Wang, Yifan Gao, Xiusi Chen, Haoming Jiang 等ICML 2024 · 被引用 52 次
