Synthetic Data is an Elegant GIFT for Continual Vision-Language Models
Bin Wu, Wuxuan Shi, Jinqiao Wang, Mang Ye
摘要
Pre-trained Vision-Language Models (VLMs) require Continual Learning (CL) to efficiently update their knowledge and adapt to various downstream tasks without retraining from scratch. However, for VLMs, in addition to the loss of knowledge previously learned from downstream tasks, pre-training knowledge is also corrupted during continual fine-tuning. This issue is exacerbated by the unavailability of original pre-training data, leaving VLM's generalization ability degrading. In this paper, we propose GIFT, a novel continual fine-tuning approach that utilizes synthetic data to overcome catastrophic forgetting in VLMs. Taking advantage of recent advances in text-to-image synthesis, we employ a pre-trained diffusion model to recreate both pretraining and learned downstream task data. In this way, the VLM can revisit previous knowledge through distillation on matching diffusion-generated images and corresponding text prompts. Leveraging the broad distribution and high alignment between synthetic image-text pairs in VLM's feature space, we propose a contrastive distillation loss along with an image-text alignment constraint. To further combat in-distribution overfitting and enhance distillation performance with limited amount of generated data, we incorporate adaptive weight consolidation, utilizing Fisher information from these synthetic image-text pairs and achieving a better stability-plasticity balance. Extensive experiments demonstrate that our method consistently outperforms previous state-of-the-art approaches across various settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- KeepLoRA: Continual Learning with Residual Gradient AdaptationMao-Lin Luo, Zi-Hao Zhou, Yi-Lin Zhang, Yuanyu Wan 等ICLR 2026 · 被引用 23 次
- Affordance-First Decomposition for Continual Learning in Video–Language UnderstandingMengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen 等CVPR 2026 · 被引用 7 次
- Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation ModelsZizhi Chen, Yizhen Gao, Minghao Han, Yizhou Liu 等CVPR 2026 · 被引用 3 次
- Spectral Imbalance Causes Forgetting in Low-Rank Continual AdaptationHao Gu, Mao-Lin Luo, Zi-Hao Zhou, Han-Chen Zhang 等ICML 2026 · 被引用 3 次
- Dance Across Shifts: Forward-Facilitation Continual Test-Time Adaptation through Dynamic Style BridgingZhilin Zhu, Yabin Wang, Zhiheng Ma, Yaguang Song 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
相关 Paper
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query TokensJian Yang, Dacheng Yin, Xiaoxuan He, Yong Li 等CVPR 2026 · 被引用 1 次
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin 等ICCV 2023 · 被引用 133 次
- How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?Jiahua Dong, Wenqi Liang, Hongliu Li, Duzhen Zhang 等NeurIPS 2024 · 被引用 42 次
- Embracing Language Inclusivity and Diversity in CLIP through Continual Language LearningBang Yang, Yong Dai, Xuxin Cheng, Yaowei Li 等AAAI 2024 · 被引用 9 次
- Mitigating Catastrophic Forgetting in Large Language Models with Self-Synthesized RehearsalJianheng Huang, Leyang Cui, Ante Wang, Chengyi Yang 等ACL 2024 · 被引用 13 次
