Rote Learning Considered Useful: Generalizing over Memorized Data in LLMs
Qinyuan Wu, Soumi Das, Mahsa Amani, Bishwamittra Ghosh, Mohammad Aflah Khan, Krishna P. Gummadi, Muhammad Bilal Zafar
摘要
Rote learning is a memorization technique based on repetition. Many researchers argue that rote learning hinders generalization because it encourages verbatim memorization rather than deeper understanding. This concern extends even to factual knowledge, which inevitably requires a certain degree of memorization. In this work, we challenge this view and demonstrate that large language models (LLMs) can, in fact, generalize over rote memorized data. We introduce a two-phase "memorize-then-generalize" framework, where the model first rote memorizes factual subject-object associations using a synthetic semantically meaningless key token and then learns to generalize by fine-tuning on a small set of semantically meaningful prompts. Extensive experiments over 8 LLMs show that the models can reinterpret rote memorized data through the semantically meaningful prompts, as evidenced by the emergence of structured, semantically aligned latent representations between the key token and the semantically meaningful prompts. This surprising finding opens the door to both effective and efficient knowledge injection as well as possible risks of repurposing the memorized data for malicious usage. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Memorization, Emergence, and Explaining Reversal Failures: A Controlled Study of Relational Semantics in LLMsYihua Zhu, Qianying Liu, Jiaxin Wang, Fei Cheng 等ACL 2026 · 被引用 1 次
- Beyond Temperature: Hyperfitting as a Late-Stage Geometric ExpansionMeimingwei Li, Yuanhao Ding, Esteban Garces Arias, Christian HeumannICML 2026
它引用的顶会 Paper27
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"Lukas Berglund, Meg Tong, Maximilian Kaufmann, Mikita Balesni 等ICLR 2024 · 被引用 462 次
相关 Paper
- Unlearning or Obfuscating? Jogging the Memory of Unlearned LLMs via Benign RelearningShengyuan Hu, Yiwei Fu, Steven Z. Wu, Virginia SmithICLR 2025
- Demystifying Verbatim Memorization in Large Language ModelsJing Huang, Diyi Yang, Christopher PottsEMNLP 2024 · 被引用 3 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Meaningful Learning: Enhancing Abstract Reasoning in Large Language Models via Generic Fact GuidanceKai Xiong, Xiao Ding, Ting Liu, Bing Qin 等NeurIPS 2024
- Memorization Sinks: Isolating Memorization during LLM TrainingGaurav Rohit Ghosal, Pratyush Maini, Aditi RaghunathanICML 2025
