Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping
Kaustubh Vijaykumar Pethkar, Ziyang Xiong, Zuofeng Shang, Yingcong Li
Abstract
Continual incorporation of new knowledge is essential for the long-term evolution of large language models (LLMs). Existing approaches typically rely on parameter-update algorithms to mitigate catastrophic forgetting, yet they suffer from fundamental limitations: 1) forgetting is unavoidable as the amount of newly injected knowledge grows; and 2) model updates are often irreversible. As modern LLMs become increasingly expressive, it is natural to question whether large-scale weight updates are necessary for acquiring a small amount of new knowledge. In this work, we propose a principled framework that models autoregressive language generation as a Markov process over tokens, where model memory is represented by a Markov transition matrix. Under this formulation, incorporating new knowledge/tokens corresponds to extending the state space, and preserving existing transitions guarantees retention of previously learned knowledge. We then prove a sample complexity bound for incorporating new tokens via a token-to-dictionary mapping strategy. In particular, for learning the transition behavior of each new token, the required number of samples scales linearly with the number of existing tokens it is mapped to. To realize this mapping, we propose an embedding-tuning algorithm that requires minimal parameter updates and induces zero forgetting. Experimental results further demonstrate the effectiveness of our method and validate our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9bc3fb4d-1e10-4b4b-a3ae-b20c25089deaBuilds on16
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Training Chain-of-Thought via Latent-Variable InferenceMatthew Douglas Hoffman, Du Phan, David Dohan, Sholto Douglas et al.NeurIPS 2023 · 74 citations
Related papers
- Train-Attention: Meta-Learning Where to Focus in Continual Knowledge LearningYeongbin Seo, Dongha Lee, Jinyoung YeoNeurIPS 2024 · 5 citations
- Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to AlgorithmHaoyu Wang, yifan shang, Zhongxiang Sun, Weijie Yu et al.ICML 2026
- Progressive Prompts: Continual Learning for Language ModelsAnastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa et al.ICLR 2023 · 15 citations
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 36 citations
- Multimodal Continual Instruction Tuning with Dynamic Gradient GuidanceSongze Li, Mingyu Gao, Tonghua Su, Xu-Yao Zhang et al.CVPR 2026 · 6 citations
