Memory Layers at Scale
Vincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, Gargi Ghosh
Abstract
Memory layers use a trainable key-value lookup mechanism to add extra parameters to a model without increasing FLOPs. Conceptually, sparsely activated memory layers complement compute-heavy dense feed-forward layers, providing dedicated capacity to store and retrieve information cheaply. This work takes memory layers beyond proof-of-concept, proving their utility at contemporary scale. On downstream tasks, language models augmented with our improved memory layer outperform dense models with more than twice the computation budget, as well as mixture-of-expert models when matched for both compute and parameters. We find gains are especially pronounced for factual tasks. We provide a fully parallelizable memory layer implementation, demonstrating scaling laws with up to 128B memory parameters, pretrained to 1 trillion tokens, comparing to base models with up to 8B parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2e44454a-5a10-47bb-985b-2b2aff5cca32Cited by top-tier papers14
- Titans: Learning to Memorize at Test TimeAli Behrouz, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 368 citations
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language ModelsXin Cheng, Wangding Zeng, Damai Dai, Qinyu Chen et al.ACL 2026 · 57 citations
- SonicMoE: Accelerating MoE with IO and Tile-aware OptimizationsWentao Guo, Mayank Mishra, Xinle Cheng, Ion Stoica et al.ICLR 2026 · 22 citations
- Scaling Embedding Layers in Language ModelsDa Yu, Edith Cohen, Badih Ghazi, Yangsibo Huang et al.NeurIPS 2025 · 21 citations
- Is the Reversal Curse a Binding Problem? Uncovering Limitations of Transformers from a Basic Generalization FailureBoshi Wang, Huan SunICLR 2026 · 16 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
Related papers
- UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context LearningZihao Huang, Yu Bao, Qiyang Min, Siyan Chen et al.ICLR 2026 · 6 citations
- Pretraining with hierarchical memories: separating long-tail and common knowledgeHadi Pouransari, David Grangier, C Thomas, Michael Kirchhof et al.ICLR 2026 · 11 citations
- Ultra-Sparse Memory NetworkZihao Huang, Qiyang Min, Hongzhi Huang, Yutao Zeng et al.ICLR 2025
- Towards A Unified View of Sparse Feed-Forward Network in Pretraining Large Language ModelZeyu Liu, Tim Dettmers, Xi Lin, Veselin Stoyanov et al.EMNLP 2023 · 3 citations
- Fine-tuning Image Transformers using Learnable MemoryMark Sandler, Andrey Zhmoginov, Max Vladymyrov, Andrew JacksonCVPR 2022 · 51 citations
