Mitigating the Evolving Semantic Entanglement in Continual Learning of Vision-Language Models
Yiliang Zhu, Dayan Wu, Qinghang Su, Zexian Yang, Zheng Lin, Weiping Wang
Abstract
Multi-domain Task Incremental Learning (MTIL) aims to continuously acquire knowledge from diverse domains while maintaining generalization capability. Recent works have demonstrated promising results by leveraging Vision-Language Models (VLMs) for continual learning. However, we identify a critical issue within this paradigm, termed the Evolving Semantic Entanglement. Specifically, VLMs tend to produce highly similar text features for semantically related categories, resulting in hard feature alignment and interference between related categories. This problem becomes increasingly pronounced as the category space expands. In this paper, we present a novel Dual-granularity Prompt Learning (DuPLe) framework to address this challenge. Our approach enhances text feature discriminability by leveraging complementary two-level prompts: category-level global prompts for holistic semantic concepts and attribute-level local prompts for fine-grained visual patterns. We further apply a Task Assignment-free Inference strategy that eliminates explicit task identification, simplifying the inference process and enabling extension to unseen categories. Extensive experiments on 11 diverse domains under MTIL and X-TAIL settings demonstrate that our method significantly mitigates the entanglement issue and outperforms previous state-of-the-art approaches.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 336da0b9-d736-4ca0-bdfa-2de6136d79d9Related papers
- Semantic-Guided Global-Local Collaborative Prompt Learning for Few-Shot Class Incremental Learningyongxin yan, Weisen Chen, Xingye Chen, Yuanjie Shao et al.CVPR 2026
- Learning Yourself: Class-Incremental Semantic Segmentation with Language-Inspired Bootstrapped DisentanglementRuitao Wu, Yifan Zhao, Jia LiICCV 2025
- Weighted Multi-Prompt Learning with Description-free Large Language Model DistillationSua Lee, Kyubum Shin, Jung Ho ParkICLR 2025
- DePro: Domain Ensemble using Decoupled Prompts for Universal Cross-Domain RetrievalKaixiang Chen, Pengfei Fang, Hui XueSIGIR 2025 · 2 citations
- Dual Prototype Evolving for Test-Time Generalization of Vision-Language ModelsCe Zhang, Simon Stepputtis, Katia P. Sycara, Yaqi XieNeurIPS 2024 · 57 citations
