Interweaving Memories of a Siamese Large Language Model
Xin Song, Zhikai Xue, Guoxiu He, Jiawei Liu, Wei Lu
摘要
Parameter-efficient fine-tuning (PEFT) methods optimize large language models (LLMs) by modifying or introducing a small number of parameters to enhance alignment with downstream tasks. However, they can result in catastrophic forgetting, where LLMs prioritize new knowledge at the expense of comprehensive world knowledge. A promising approach to mitigate this issue is to recall prior memories based on the original knowledge. To this end, we propose a model-agnostic PEFT framework, IMSM, which Interweaves Memories of a Siamese Large Language Model. Specifically, our siamese LLM is equipped with an existing PEFT method. Given an incoming query, it generates two distinct memories based on the pre-trained and fine-tuned parameters. IMSM then incorporates an interweaving mechanism that regulates the contributions of both original and enhanced memories when generating the next token. This framework is theoretically applicable to all open-source LLMs and existing PEFT methods. We conduct extensive experiments across various benchmark datasets, evaluating the performance of popular open-source LLMs using the proposed IMSM, in comparison to both classical and leading PEFT methods. Our findings indicate that IMSM maintains comparable time and space efficiency to backbone PEFT methods while significantly improving performance and effectively mitigating catastrophic forgetting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil 等ICLR 2024 · 被引用 1,798 次
- Few-Shot Parameter-Efficient Fine-Tuning is Better and Cheaper than In-Context LearningHaokun Liu, Derek Tam, Mohammed Muqeeth, Jay Mohta 等NeurIPS 2022 · 被引用 1,483 次
- DoRA: Weight-Decomposed Low-Rank AdaptationShih-Yang Liu, Chien-Yi Wang, Hongxu Yin, Pavlo Molchanov 等ICML 2024 · 被引用 820 次
相关 Paper
- PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from AttentionHaonan Wang, Brian K Chen, Siquan Li, Liang Xinhe 等ICLR 2026 · 被引用 5 次
- LoKI: Low-Damage Knowledge Implanting of Large Language ModelsRunyu Wang, Peng Ping, Zhengyu Guo, Xiaoye Zhang 等AAAI 2026 · 被引用 3 次
- HFT: Half Fine-Tuning for Large Language ModelsTingfeng Hui, Zhenyu Zhang, Shuohuan Wang, Weiran Xu 等ACL 2025
- The Inter-Intra Modal Measure: A Predictive Lens on Fine-Tuning Outcomes in Vision-Language ModelsLaura Niss, Kevin Vogt-Lowell, Theodoros TsiligkaridisICCV 2025 · 被引用 1 次
- Self-Distillation Bridges Distribution Gap in Language Model Fine-TuningZhaorui Yang, Tianyu Pang, Haozhe Feng, Han Wang 等ACL 2024
