Continual Learners are Incremental Model Generalizers
Jaehong Yoon, Sung Ju Hwang, Yue Cao
摘要
Motivated by the efficiency and rapid convergence of pre-trained models for solving downstream tasks, this paper extensively studies the impact of Continual Learning (CL) models as pre-trainers. In both supervised and unsupervised CL, we find that the transfer quality of the representation often increases gradually without noticeable degradation in fine-tuning performance. This is because CL models can learn improved task-general features when easily forgetting task-specific knowledge. Based on this observation, we suggest a new unsupervised CL framework with masked modeling, which aims to capture fluent task-generic representation during training. Furthermore, we propose a new fine-tuning scheme, GLobal Attention Discretization (GLAD), that preserves rich task-generic representation during solving downstream tasks. The model fine-tuned with GLAD achieves competitive performance and can also be used as a good pre-trained model itself. We believe this paper breaks the barriers between pre-training and fine-tuning steps and leads to a sustainable learning framework in which the continual learner incrementally improves model generalization, yielding better transfer to unseen tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringFeifei Zhang, Zhihao Wang, Xi Zhang, Changsheng XuICCV 2025 · 被引用 3 次
- STELLA: Continual Audio-Video Pre-training with SpatioTemporal Localized AlignmentJaewoo Lee, Jaehong Yoon, Wonjae Kim, Yunji Kim 等ICML 2024 · 被引用 2 次
它引用的顶会 Paper23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Continual Pre-training of Language ModelsZixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi 等ICLR 2023 · 被引用 15 次
- Beyond Sharpness: A Flatness Decomposition Framework for Efficient Continual LearningYanan Chen, Tieliang Gong, Yunjiao Zhang, Wen WenAAAI 2026
- Enhancing Visual Continual Learning with Language-Guided SupervisionBolin Ni, Hongbo Zhao, Chenghao Zhang, Ke Hu 等CVPR 2024 · 被引用 8 次
- Detect, Decide, Unlearn: A Transfer-Aware Framework for Continual LearningYiwen Wang, Diana Benavides-Prado, Yun Sing KohICLR 2026
- GLID: Pre-training a Generalist Encoder-Decoder Vision ModelJihao Liu, Jinliang Zheng, Yu Liu, Hongsheng LiCVPR 2024 · 被引用 5 次
