Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
Xinyi Wang, Antonis Antoniades, Yanai Elazar, Alfonso Amayuelas, Alon Albalak, Kexun Zhang, William Yang Wang
摘要
The impressive capabilities of large language models (LLMs) have sparked debate over whether these models genuinely generalize to unseen tasks or predominantly rely on memorizing vast amounts of pretraining data. To explore this issue, we introduce an extended concept of memorization, distributional memorization, which measures the correlation between the LLM output probabilities and the pretraining data frequency. To effectively capture task-specific pretraining data frequency, we propose a novel task-gram language model, which is built by counting the co-occurrence of semantically related -gram pairs from task inputs and outputs in the pretraining corpus. Using the Pythia models trained on the Pile dataset, we evaluate four distinct tasks: machine translation, factual question answering, world knowledge understanding, and math reasoning. Our findings reveal varying levels of memorization, with the strongest effect observed in factual question answering. Furthermore, while model performance improves across all tasks as LLM size increases, only factual question answering shows an increase in memorization, whereas machine translation and reasoning tasks exhibit greater generalization, producing more novel outputs. This study demonstrates that memorization plays a larger role in simpler, knowledge-intensive tasks, while generalization is the key for harder, reasoning-based tasks, providing a scalable method for analyzing large pretraining corpora in greater depth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- Tracing the Representation Geometry of Language Models from Pretraining to Post-trainingMelody Zixuan Li, Kumar Krishna Agrawal, Arna Ghosh, Komal Kumar Teru 等NeurIPS 2025 · 被引用 38 次
- Hubble: a Model Suite to Advance the Study of LLM MemorizationJohnny Wei, Ameya Godbole, Mohammad Aflah Khan, Ryan Yixiang Wang 等ICLR 2026 · 被引用 22 次
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent ThoughtsHanwen Du, Yuxin Dong, Xia NingICLR 2026 · 被引用 20 次
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning ModelTianle Li, Jihai Zhang, Yongming Rao, Yu ChengNeurIPS 2025 · 被引用 17 次
- Attack via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMsZhixin Xie, Xurui Song, Jun LuoNeurIPS 2025 · 被引用 11 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 被引用 1,030 次
相关 Paper
- Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention LayersLei Chen, Joan Bruna, Alberto BiettiICLR 2025
- What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?Katie Kang, Amrith Setlur, Dibya Ghosh, Jacob Steinhardt 等ICML 2025
- Procedural Knowledge in Pretraining Drives Reasoning in Large Language ModelsLaura Ruis, Maximilian Mozes, Juhan Bae, Siddhartha Rao Kamalakara 等ICLR 2025
- How Do Large Language Models Acquire Factual Knowledge During Pretraining?Hoyeon Chang, Jinho Park, Seonghyeon Ye, Sohee Yang 等NeurIPS 2024 · 被引用 124 次
- Shared Path: Unraveling Memorization in Multilingual LLMs through Language SimilaritiesXiaoyu Luo, Yiyi Chen, Johannes Bjerva, Qiongxiu LiEMNLP 2025 · 被引用 1 次
