Pretraining with hierarchical memories: separating long-tail and common knowledge
Hadi Pouransari, David Grangier, C Thomas, Michael Kirchhof, Oncel Tuzel
Abstract
Apple he impressive performance gains of modern language models currently rely on scaling parameters: larger models store more world knowledge and reason better. Yet compressing all world knowledge into parameters is unnecessary, as only a fraction is used per prompt, and impractical for edge devices with limited inference-time memory and compute. We address this shortcoming by a memory-augmented architecture and a pretraining strategy aligned with existing hardware paradigms. We introduce small language models that access large hierarchical parametric memory banks encoding world knowledge. During pretraining and inference, we fetch a small, context-dependent memory block and add it to the model. Our pretraining learns to store long-tail world knowledge in the memory parameters, while the small language model acts as an anchor capturing common knowledge and general reasoning abilities. Through trillion-token-scale experiments, we show significant gains: a 160M-parameters model augmented with an 18M-parameters memory fetched from a 4.6B memory bank obtains comparable performance to a regular model with more than 2× the parameters. Through extensive experiments, we study the optimal type and size of parametric memories in transformers, scaling them to over 21B parameters. We find that our proposed hierarchical feed-forward memories work robustly across transformer architectures, whether added during pretraining or post-hoc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42498e35-8773-414a-b25b-11ab11b0de87Cited by top-tier papers3
- Procedural Pretraining: Warming Up Language Models with Abstract DataLiangze Jiang, Zachary Shinnick, Anton Hengel, Hemanth Saratchandran et al.ICML 2026 · 6 citations
- Optimal Splitting of Language Models from Mixtures to Specialized DomainsSkyler Seto, Pierre Ablin, Anastasiia Filippova, Jiayuan Ye et al.ICML 2026 · 2 citations
- Cram Less to Fit More: Training Data Pruning Improves Memorization of FactsJiayuan Ye, Vitaly Feldman, Kunal TalwarICML 2026 · 1 citation
Builds on22
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
Related papers
- Memory Layers at ScaleVincent-Pierre Berges, Barlas Oguz, Daniel Haziza, Wen-tau Yih et al.ICML 2025
- Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language ModelsXiaoman Pan, Wenlin Yao, Hongming Zhang, Dian Yu et al.ICLR 2023 · 2 citations
- Understanding LoRA as Knowledge Memory: An Empirical AnalysisSeungju Back, Dongwoo Lee, Naun Kang, Taehee Lee et al.ICML 2026 · 10 citations
- Beyond Attention: Breaking the Limits of Transformer Context Length with Recurrent MemoryAydar Bulatov, Yuri Kuratov, Yermek Kapushev, Mikhail BurtsevAAAI 2024 · 22 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
