Subspace Alignment for CLIP-based Continual Learning via Canonical Correlation Analysis
Huan Zhang, Shuyu Dong, Yujin Zheng, Dingwen Wang, Shenghua Fan, Fan Lyu
摘要
Recent advances in CLIP-based continual learning have shown the potential of leveraging pre-trained vision-language models for sequential tasks. However, existing methods overlook a key problem we call Asymmetric Drift. In unimodal CLIP-based continual learning, the visual branch undergoes stronger adaptation because the visual distribution shifts significantly, whereas the text branch remains relatively stable due to the low variance of textual prompts. This imbalance increases the modality distance and degrades cross-modal alignment over time. To address this issue, we propose CCA-CL, a framework that accumulates visual-textual covariance statistics across tasks and solves Canonical Correlation Analysis to compute a shared subspace. In this subspace, the distance between visual and textual features is minimized, enabling better alignment without modifying CLIP parameters. This also makes our method naturally compatible with exemplar-free CL settings. To further capture nonlinear relationships that linear Canonical Correlation Analysis is hard to model, we introduce Random Fourier Projection as an extension. Experimental results demonstrate that CCA-CL effectively mitigates the asymmetric drift problem and achieves state-ofthe-art performance on several benchmarks. We release the code at https://github.com/zhwhu/CCA-CL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath 等ICCV 2021 · 被引用 2,294 次
- Learning to Prompt for Continual LearningZifeng Wang, Zizhao Zhang, Chen-Yu Lee, Han Zhang 等CVPR 2022 · 被引用 635 次
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 被引用 419 次
- From Canonical Correlation Analysis to Self-supervised Graph Neural NetworksHengrui Zhang, Qitian Wu, Junchi Yan, David Wipf 等NeurIPS 2021 · 被引用 319 次
- Using Hindsight to Anchor Past Knowledge in Continual LearningArslan Chaudhry, Albert Gordo, Puneet K. Dokania, Philip H. S. Torr 等AAAI 2021 · 被引用 279 次
相关 Paper
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin 等ICCV 2023 · 被引用 133 次
- Dynamic Multi-Layer Null Space Projection for Vision-Language Continual LearningBorui Kang, Lei Wang, Zhiping Wu, Tao Feng 等ICCV 2025 · 被引用 4 次
- Pi-CCA: Prompt-Invariant CCA Certificates for Replay-Free Continual Multimodal LearningJiayu Zhang, Chuangxin Zhao, Canran Xiao, Ruibo Duan 等ICLR 2026
- Continual Vision-Language Representation Learning with Off-Diagonal InformationZixuan Ni, Longhui Wei, Siliang Tang, Yueting Zhuang 等ICML 2023 · 被引用 40 次
