CoMem: Compositional Concept-Graph Memory for Vision-Language Adaptation
Heng Zhou, Jing Tang, Jusheng Zhang, Yanshu Li, Canran Xiao, Liwei Hou, Zong Ke, Jiawei Yao
Abstract
Continual vision–language learning is crucial for multimodal tasks such as image–text retrieval, visual question answering, and grounded reasoning in dynamic environments, yet deployed systems must learn from non-stationary streams under strict privacy and memory budgets, where naïve finetuning forgets and harms transfer. We aim to sustain stable yet plastic capability in this setting without storing raw data, enabling reuse and recombination across domains and tasks. We present CoMem, a framework that treats compositional structure as the unit of memory and rehearsal: it incrementally organizes knowledge into a compact graph of concepts and relations and rehearses directly in feature space by conditioning practice signals on sampled subgraphs. A lightweight compositional consistency objective keeps part–whole predictions coherent, while teacher-informed, uncertainty-aware filtering limits off-manifold drift. Across cross-domain retrieval, structured concept learning, and continual multimodal VQA, CoMem achieves state-of-the-art retention and transfer alongside consistent gains on SVLC and VQACL/CLOVE under matched memory and parameter budgets. By casting structure as memory and rehearsing where learning happens (feature space), CoMem provides a privacy-friendly and testable paradigm for reliable continual adaptation without raw exemplars.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38cd294b-3902-4118-9f09-d22f2cf365d1Builds on30
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Follow Your Pose: Pose-Guided Text-to-Video Generation Using Pose-Free VideosYue Ma, Yingqing He, Xiaodong Cun, Xintao Wang et al.AAAI 2024 · 318 citations
- BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual QuestionsWenbo Hu, Yifan Xu, Yi Li, Weiyue Li et al.AAAI 2024 · 209 citations
- Preventing Zero-Shot Transfer Degradation in Continual Learning of Vision-Language ModelsZangwei Zheng, Mingyuan Ma, Kai Wang, Ziheng Qin et al.ICCV 2023 · 133 citations
- Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts AdaptersJiazuo Yu, Yunzhi Zhuge, Lu Zhang, Ping Hu et al.CVPR 2024 · 80 citations
Related papers
- Towards General Continuous Memory for Vision-Language ModelsWenyi Wu, Zixuan Song, Kun Zhou, Yifei Shao et al.NeurIPS 2025 · 19 citations
- Reversible Primitive-Composition Alignment for Continual Vision-Language LearningCanran Xiao, Tianxiang Xu, SiYuan Ma, Yiyang Jiang et al.ICLR 2026
- MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question AnsweringZhifei Li, Yiran Wang, Chenyi Xiong, Yujing Xia et al.AAAI 2026
- VQACL: A Novel Visual Question Answering Continual Learning SettingXi Zhang, Feifei Zhang, Changsheng XuCVPR 2023
- Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question AnsweringImad Eddine Marouf, Enzo Tartaglione, Stéphane Lathuilière, Joost van de WeijerICCV 2025 · 4 citations
