Continual Learning with Global Alignment
Xueying Bai, Jinghuan Shang, Yifan Sun, Niranjan Balasubramanian
摘要
Continual learning aims to sequentially learn new tasks without forgetting previous tasks' knowledge (catastrophic forgetting). One factor that can cause forgetting is the interference between the gradients on losses from different tasks. When the gradients on the current task's loss are in opposing directions to those on previous tasks' losses, updating the model for the current task may cause performance degradation on previous tasks. In this paper, we first identify causes of the above interference, and hypothesize that correlations between data representations are a key factor of interference. We then propose a method for promoting appropriate correlations between arbitrary tasks' data representations (i.e., global alignment) in individual task learning. Specifically, we learn the data representation as a taskspecific composition of pre-trained token representations shared across all tasks. Then the correlations between different tasks' data representations are grounded by correlations between pre-trained token representations. We explore different ways to learn such compositions. Without experience replay, our model achieves SOTA performance in continual learning tasks. It also achieves advanced classincremental performance through task-incremental training. The code is available at: https://github.com/StonyBrookNLP/global-alignment .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
相关 Paper
- AnaCP: Toward Upper-Bound Continual Learning via Analytic Contrastive ProjectionSaleh Momeni, Changnan Xiao, Bing LiuNeurIPS 2025 · 被引用 8 次
- STAR: Stability-Inducing Weight Perturbation for Continual LearningMasih Eskandar, Tooba Imtiaz, Davin Hill, Zifeng Wang 等ICLR 2025
- Prototype-Sample Relation Distillation: Towards Replay-Free Continual LearningNader Asadi, MohammadReza Davari, Sudhir P. Mudur, Rahaf Aljundi 等ICML 2023 · 被引用 61 次
- CafeBoost: Causal Feature Boost to Eliminate Task-Induced Bias for Class Incremental LearningBenliu Qiu, Hongliang Li, Haitao Wen, Heqian Qiu 等CVPR 2023
- Layerwise Optimization by Gradient Decomposition for Continual LearningShixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu 等CVPR 2021
