Learning from Itself: Mining Internal Knowledge from Vision Language Models for Continual Learning
Yizheng Gong, Siyue Yu, Waleed Al-Nuaimy, Jimin Xiao
摘要
Vision-language models like CLIP excel at zero-shot recognition but struggle with continual learning due to two critical issues: (1) severe distribution gap between pretraining captions and post-training class names, and (2) performance mismatch between vision-only and dual-encoder approaches—vision-only methods achieve 20% higher accuracy on fine-grained tasks while CLIP dominates on natural images. We propose Learning from Itself (LfI), which mines CLIP's internal knowledge to address both challenges. First, we generate pseudo-captions by optimizing learnable tokens to minimize CLIP's contrastive loss, creating auxiliary training signals that bridge the pretraining-finetuning distribution gap without external models. Second, we introduce adaptive mutual distillation that dynamically weights knowledge transfer between CLIP's text encoder and a temporary vision classifier based on their instantaneous performance—stronger branches teach more, weaker ones learn more. At inference, only the original CLIP architecture is used, having absorbed discriminative knowledge from both branches. LfI achieves state-of-the-art results across multiple continual learning benchmarks, demonstrating that CLIP can effectively teach itself to continually learn new tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
相关 Paper
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin 等ICCV 2023 · 被引用 133 次
- Learning Mask-aware CLIP Representations for Zero-Shot SegmentationSiyu Jiao, Yunchao Wei, Yaowei Wang, Yao Zhao 等NeurIPS 2023 · 被引用 88 次
- C-CLIP: Multimodal Continual Learning for Vision-Language ModelWenzhuo Liu, Fei Zhu, Longhui Wei, Qi TianICLR 2025
- Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual LearningLinlan Huang, Xusheng Cao, Haori Lu, Yifan Meng 等ICCV 2025 · 被引用 12 次
- Towards Robustness Prompt Tuning with Fully Test-Time Adaptation for CLIP's Zero-Shot GeneralizationRan Wang, Hua Zuo, Zhen Fang, Jie LuACM MM 2024 · 被引用 7 次
