Plug-and-Play Compositionality for Boosting Continual Learning with Foundation Models
Weiduo Liao, Fei Han, Hisao Ishibuchi, Qingfu Zhang, Ying Wei
Abstract
Vision learners often struggle with catastrophic forgetting due to their reliance on class recognition by comparison, rather than understanding classes as compositions of representative concepts. This limitation is prevalent even in state-of-the-art continual learners with foundation models and worsens when current tasks contain few classes. Inspired by the recent success of concept-level understanding in mitigating forgetting, we design a universal framework CompSLOT to guide concept learning across diverse continual learners. Leveraging the progress of object-centric learning in parsing semantically meaningful slots from images, we tackle the challenge of learning slot extraction from ImageNet-pretrained vision transformers by analyzing meaningful concept properties. We further introduce a primitive selection and aggregation mechanism to harness concept-level image understanding. Additionally, we propose a method-agnostic self-supervision approach to distill sample-wise concept-based similarity information into the classifier, reducing reliance on incorrect or partial concepts for classification. Experiments show CompSLOT significantly enhances various continual learners and provides a universal concept-level module for the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d20a8b1-3915-443e-895d-1cd9b20940a4Builds on46
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Attention Retention for Continual Learning with Vision TransformersYue Lu, Xiangyu Zhou, Shizhou Zhang, Yinghui Xing et al.AAAI 2026
- Continual Learning with Lifelong Vision TransformerZhen Wang, Liu Liu, Yiqun Duan, Yajing Kong et al.CVPR 2022 · 63 citations
- MUFASA: A Multi-Layer Framework for Slot AttentionSebastian Bock, Leonie Schüßler, Krishnakant Singh, Simone Schaub-Meyer et al.CVPR 2026 · 1 citation
- Convolutional Prompting meets Language Models for Continual LearningAnurag Roy, Riddhiman Moulick, Vinay Kumar Verma, Saptarshi Ghosh et al.CVPR 2024 · 15 citations
- CoMFormer: Continual Learning in Semantic and Panoptic SegmentationFabio Cermelli, Matthieu Cord, Arthur DouillardCVPR 2023
