ConStruct-VL: Data-Free Continual Structured VL Concepts Learning
James Seale Smith, Paola Cascante-Bonilla, Assaf Arbelle, Donghyun Kim, Rameswar Panda, David D. Cox, Diyi Yang, Zsolt Kira, Rogério Feris, Leonid Karlinsky
Abstract
Recently, large-scale pre-trained Vision-and-Language (VL) foundation models have demonstrated remarkable capabilities in many zero-shot downstream tasks, achieving competitive results for recognizing objects defined by as little as short text prompts. However, it has also been shown that VL models are still brittle in Structured VL Concept (SVLC) reasoning, such as the ability to recognize object attributes, states, and inter-object relations. This leads to reasoning mistakes, which need to be corrected as they occur by teaching VL models the missing SVLC skills; often this must be done using private data where the issue was found, which naturally leads to a data-free continual (no task-id) VL learning setting. In this work, we introduce the first Continual Data-Free Structured VL Concepts Learning (ConStruct-VL) benchmark 1 and show it is challenging for many existing data-free CL strategies. We, therefore, propose a data-free method comprised of a new approach of Adversarial Pseudo-Replay (APR) which generates adversarial reminders of past tasks from past task models. To use this method efficiently, we also propose a continual parameter-efficient Layered-LoRA (LaLo) neural architecture allowing no-memory-cost access to all past models at train time. We show this approach outperforms all data-free methods by as much as ∼ 7% while even matching some levels of experience-replay (prohibitive for applications where data-privacy must be preserved). * This work is supported by the Defense Advanced Research Projects Agency (DARPA) Contract No. FA8750-19-C-1001. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of DARPA. † Equal contribution 1 Our code is publicly available at https : / / github . com / jamessealesmith/ConStruct-VL Understanding Spatial relations Understanding Colors Understanding Action relations Understanding Object states Reminding using Adversarial Pseudo-Replay (APR) Efficient access to past models with Layered LoRA (LaLo) architecture Continual Learning of Structured V&L Concepts
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig et al.NeurIPS 2023 · 93 citations
- Going Beyond Nouns With Vision & Language Models Using Synthetic DataPaola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh et al.ICCV 2023 · 49 citations
- Gated Integration of Low-Rank Adaptation for Continual Learning of Large Language ModelsYan-Shuo Liang, Jia-Rui Chen, Wu-Jun LiNeurIPS 2025 · 15 citations
- Stabilizing Zero-Shot Prediction: A Novel Antidote to Forgetting in Continual Vision-Language TasksZijian Gao, Xingxing Zhang, Kele Xu, Xinjun Mao et al.NeurIPS 2024 · 11 citations
- Affordance-First Decomposition for Continual Learning in Video–Language UnderstandingMengzhu xu, Hanzhi Liu, Ningkang Peng, qianyu Chen et al.CVPR 2026 · 7 citations
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- Pretrained Vision-Language-Action Models are Surprisingly Resistant to Forgetting in Continual LearningHuihan Liu, Changyeon Kim, Bo Liu, Minghuan Liu et al.ICML 2026 · 14 citations
- Reversible Primitive-Composition Alignment for Continual Vision-Language LearningCanran Xiao, Tianxiang Xu, SiYuan Ma, Yiyang Jiang et al.ICLR 2026
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic ForgettingAsher J. Hancock, Xindi Wu, Lihan Zha, Olga Russakovsky et al.ICLR 2026 · 58 citations
- LoRA Recycle: Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAsZixuan Hu, Yongxian Wei, Li Shen, Chun Yuan et al.CVPR 2025
