Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm
Haoyu Wang, yifan shang, Zhongxiang Sun, Weijie Yu, Xiao Zhang, Jun Xu
摘要
Continual Pre-Training (CPT) is essential for enabling Language Models (LMs) to integrate new factual knowledge without erasing old. While classical CPT techniques like data replay have become the standard paradigm, the mechanisms underlying how LMs acquire and retain facts over time, termed as continual Factual Knowledge Acquisition (cFKA), remain unclear. In this work, we present a theoretical framework that characterizes the training dynamics of cFKA using a single-layer Transformer with linear attention, offering a unified explanation for the behavior of popular CPT methods. Our analysis reveals that regularization-based methods merely adjust the convergence rate of parameters without altering the inherent forgetting tendency, whereas data replay methods shift convergence dynamics and stabilize pretrained knowledge. Building on these insights, we propose a novel generative data replay approach, called Selecting Tokens via attentiOn Contribution (STOC), which identifies influential factual snippets to guide replay generation. Extensive experiments on both synthetic and real-world datasets validate our theoretical findings and demonstrate that STOC effectively enhances cFKA by mitigating catastrophic forgetting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Transformers learn to implement preconditioned gradient descent for in-context learningKwangjun Ahn, Xiang Cheng, Hadi Daneshmand, Suvrit SraNeurIPS 2023 · 被引用 324 次
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 被引用 258 次
相关 Paper
- Progressive Prompts: Continual Learning for Language ModelsAnastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa 等ICLR 2023 · 被引用 15 次
- SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language ModelsJinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao 等EMNLP 2024 · 被引用 4 次
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-TrainingSong Lai, Haohan Zhao, Rong Feng, Changyi Ma 等ICML 2026 · 被引用 46 次
- Train-Attention: Meta-Learning Where to Focus in Continual Knowledge LearningYeongbin Seo, Dongha Lee, Jinyoung YeoNeurIPS 2024 · 被引用 5 次
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
