Meta-attention for ViT-backed Continual Learning
Mengqi Xue, Haofei Zhang, Jie Song, Mingli Song
Abstract
Continual learning is a longstanding research topic due to its crucial role in tackling continually arriving tasks. Up to now, the study of continual learning in computer vision is mainly restricted to convolutional neural networks (CNNs). However, recently there is a tendency that the newly emerging vision transformers (ViTs) are gradually dominating the field of computer vision, which leaves CNN-based continual learning lagging behind as they can suffer from severe performance degradation if straightforwardly applied to ViTs. In this paper, we study ViT-backed continual learning to strive for higher performance riding on recent advances of ViTs. Inspired by mask-based continual learning methods in CNNs, where a mask is learned per task to adapt the pre-trained ViT to the new task, we propose MEta-ATtention (MEAT), i.e., attention to self-attention, to adapt a pre-trained ViT to new tasks without sacrificing performance on already learned tasks. Unlike prior mask-based methods like Piggyback, where all parameters are associated with corresponding masks, MEAT leverages the characteristics of ViTs and only masks a portion of its parameters. It renders MEAT more efficient and effective with less overhead and higher accuracy. Extensive experiments demonstrate that MEAT exhibits significant superiority to its state-of-the-art CNN counterparts, with 4.0 ∼ 6.0% absolute boosts in accuracy. Our code has been released at https://github.com/zju-vipa/MEAT-TIL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97ea3338-b93b-4de5-8afd-0727df5e7ae9Cited by top-tier papers15
- Convolutional Prompting meets Language Models for Continual LearningAnurag Roy, Riddhiman Moulick, Vinay Kumar Verma, Saptarshi Ghosh et al.CVPR 2024 · 15 citations
- Exemplar-Free Continual Transformer with ConvolutionsAnurag Roy, Vinay Kumar Verma, Sravan Voonna, Kripabandhu Ghosh et al.ICCV 2023 · 13 citations
- Task-Free Continual Generation and Representation Learning via Dynamic Expansionable Memory ClusterFei Ye, Adrian G. BorsAAAI 2024 · 8 citations
- Overcoming Dual Drift for Continual Long-Tailed Visual Question AnsweringFeifei Zhang, Zhihao Wang, Xi Zhang, Changsheng XuICCV 2025 · 3 citations
- Learning Expandable and Adaptable Representations for Continual LearningRuilong Yu, Mingyan Liu, Fei Ye, Adrian G. Bors et al.NeurIPS 2025 · 3 citations
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
Related papers
- Continual Learning with Lifelong Vision TransformerZhen Wang, Liu Liu, Yiqun Duan, Yajing Kong et al.CVPR 2022 · 63 citations
- Attention Retention for Continual Learning with Vision TransformersYue Lu, Xiangyu Zhou, Shizhou Zhang, Yinghui Xing et al.AAAI 2026
- Task-Free Dynamic Sparse Vision Transformer for Continual LearningFei Ye, Adrian G. BorsAAAI 2024 · 7 citations
- Visual Prompt Tuning in Null Space for Continual LearningYue Lu, Shizhou Zhang, De Cheng, Yinghui Xing et al.NeurIPS 2024 · 42 citations
- Masked Image Residual Learning for Scaling Deeper Vision TransformersGuoxi Huang, Hongtao Fu, Adrian G. BorsNeurIPS 2023 · 10 citations
