SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
Jinghan He, Haiyun Guo, Kuan Zhu, Zihan Zhao, Ming Tang, Jinqiao Wang
摘要
Continual learning (CL) is crucial for language models to dynamically adapt to the evolving real-world demands. To mitigate the catastrophic forgetting problem in CL, data replay has been proven a simple and effective strategy, and the subsequent data-replay-based distillation can further enhance the performance. However, existing methods fail to fully exploit the knowledge embedded in models from previous tasks, resulting in the need for a relatively large number of replay samples to achieve good results. In this work, we first explore and emphasize the importance of attention weights in knowledge retention, and then propose a SElective attEntion-guided Knowledge Retention method (SEEKR) for data-efficient replay-based continual learning of large language models (LLMs). Specifically, SEEKR performs attention distillation on the selected attention heads for finer-grained knowledge retention, where the proposed forgettabilitybased and task-sensitivity-based measures are used to identify the most valuable attention heads. Experimental results on two continual learning benchmarks for LLMs demonstrate the superiority of SEEKR over the existing methods on both performance and efficiency. Explicitly, SEEKR achieves comparable or even better performance with only 1/10 of the replayed data used by other methods, and reduces the proportion of replayed data to 1%. The code is available at https: //github.com/jinghan1he/SEEKR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language ModelsYan-Shuo Liang, Jia-Rui Chen, Wu-Jun LiNeurIPS 2025 · 被引用 15 次
- Any-SSR: How Recursive Least Squares Works in Continual Learning of Large Language ModelsKai Tong, Kang Pan, Xiao Zhang, Erli Meng 等ICCV 2025 · 被引用 3 次
- FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual LearningYujie Feng, Hao Wang, Jian Li, Xu Chu 等ACL 2026 · 被引用 3 次
- PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual LearningZhiYan Hou, Haiyun Guo, Haokai Ma, Yandu Sun 等ACL 2026 · 被引用 1 次
- Recurrent Knowledge Identification and Fusion for Language Model Continual LearningYujie Feng, Xujia Wang, Zexin Lu, Shenghong Fu 等ACL 2025
它引用的顶会 Paper12
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Dark Experience for General Continual Learning: a Strong, Simple BaselinePietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati 等NeurIPS 2020 · 被引用 1,494 次
- Class-Incremental Learning by Knowledge Distillation with Adaptive Feature ConsolidationMinsoo Kang, Jaeyoo Park, Bohyung HanCVPR 2022 · 被引用 189 次
相关 Paper
- Train-Attention: Meta-Learning Where to Focus in Continual Knowledge LearningYeongbin Seo, Dongha Lee, Jinyoung YeoNeurIPS 2024 · 被引用 5 次
- SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language ModelsWeixiang Zhao, Shilong Wang, Yulin Hu, Yanyan Zhao 等ACL 2024
- Progressive Prompts: Continual Learning for Language ModelsAnastasia Razdaibiedina, Yuning Mao, Rui Hou, Madian Khabsa 等ICLR 2023 · 被引用 15 次
- Filling Memory Gaps: Enhancing Continual Semantic Parsing via SQL Syntax Variance-Guided LLMs Without Real Data ReplayRuiheng Liu, Jinyu Zhang, Yanqi Song, Yu Zhang 等AAAI 2025 · 被引用 5 次
- Learn more, but bother less: parameter efficient continual learningFuli Qiao, Mehrdad MahdaviNeurIPS 2024 · 被引用 36 次
