Training Language Models with Memory Augmentation
Zexuan Zhong, Tao Lei, Danqi Chen
摘要
Recent work has improved language models (LMs) remarkably by equipping them with a non-parametric memory component. However, most existing approaches only introduce memories at testing time or represent them using a separately trained encoder, resulting in suboptimal training of the language model. In this work, we present TRIME, a novel yet simple training approach designed for training LMs with memory augmentation. Our approach uses a training objective that directly takes inbatch examples as accessible memory. We also present new methods for memory construction and data batching, which are used for adapting to different sets of memories-local, longterm, and external memory-at testing time. We evaluate TRIME on multiple language modeling and machine translation benchmarks and show that it is able to achieve significant improvements across all the settings. Concretely, TRIME reduces the perplexity from 18.70 to 15.37 on WIKITEXT-103, by effectively leveraging a large memory set from the training corpus. Compared to standard LM training, TRIME adds negligible computational overhead and is compatible with different neural architectures, making it a versatile solution for training memory-augmented LMs. 1 * TL currently works at Google Research. The collaboration was initialized before TL joined Google. 1 Our code and pre-trained models are publicly available at https://github.com/princeton-nlp/TRIME .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper39
- Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge ConflictsJian Xie, Kai Zhang, Jiangjie Chen, Renze Lou 等ICLR 2024 · 被引用 294 次
- Focused Transformer: Contrastive Training for Context ScalingSzymon Tworkowski, Konrad Staniszewski, Mikolaj Pacek, Yuhuai Wu 等NeurIPS 2023 · 被引用 190 次
- Lift Yourself Up: Retrieval-augmented Text Generation with Self-MemoryXin Cheng, Di Luo, Xiuying Chen, Lemao Liu 等NeurIPS 2023 · 被引用 177 次
- Unlimiformer: Long-Range Transformers with Unlimited Length InputAmanda Bertsch, Uri Alon, Graham Neubig, Matthew GormleyNeurIPS 2023 · 被引用 176 次
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 被引用 152 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
相关 Paper
- Prompting Neural Machine Translation with Translation MemoriesAbudurexiti Reheman, Tao Zhou, Yingfeng Luo, Di Yang 等AAAI 2023 · 被引用 11 次
- Fast and Accurate Neural Machine Translation with Translation MemoryQiuxiang He, Guoping Huang, Qu Cui, Li Li 等ACL 2021
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof 等ICLR 2026 · 被引用 11 次
- Neural Machine Translation with Contrastive Translation MemoriesXin Cheng, Shen Gao, Lemao Liu, Dongyan Zhao 等EMNLP 2022 · 被引用 10 次
- Pre-training Limited Memory Language Models with Internal and External KnowledgeLinxi Zhao, Sofian Zalouk, Christian K. Belardi, Justin Lovelace 等ICLR 2026 · 被引用 11 次
