Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning Method
Rui Yang, Shuang Wang, Huan Zhang, Siyuan Xu, Yanhe Guo, Xiutiao Ye, Biao Hou, Licheng Jiao
Abstract
To enable machines to mimic human cognitive abilities and alleviate the catastrophic forgetting problem in cross-modal image-text retrieval (CMITR), this paper proposes a novel continual learning method, Knowledge Decomposition and Replay (KDR), which emulates the process of knowledge decomposition and replay exhibited by humans in complex and changing environments. KDR has two components: a feature Decomposition-based CMITR Model (DCM) and a cross-task Generic Knowledge Replay strategy (GKR). DCM decomposes text and image features into task-specific and generic knowledge features, mimicking the human cognitive process of knowledge decomposition. Specifically, it employs a generic knowledge features extraction module for all tasks and a task-specific module for each task with a few trainable fully connected layers. Similarly, GKR emulates the human behavior of knowledge replay by utilizing the image-text similarity matrix output from the old task model with inputting the previous samples to induce the learning of the image-text similarity matrix output from the current task model with inputting the previous samples, using knowledge distillation technology. To demonstrate the effect of KDR, we adapted a continual learning dataset Seq-COCO from MSCOCO. Extensive experiments on Seq-COCO showed that KDR reduces catastrophic forgetting and consolidates general knowledge, improving the model's learning ability in CMITR.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a85643b7-819b-4c3a-ac72-4bcba6c8f36aCited by top-tier papers3
- Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken GenerationYongqi Li, Hongru Cai, Wenjie Wang, Leigang Qu et al.SIGIR 2025 · 6 citations
- CoPE: Continual Probe-guided Expansion for Large Vision-Language ModelsZiqin Wang, Hengyuan Zhao, Qixin Sun, Kaiyou Song et al.ICML 2026
- Smart Replay: Adaptive Scheduling of Memory Rehearsal for Computational Resource-Aware Incremental LearningJianting Chen, Dianzhi Yu, Irwin KingCVPR 2026
Related papers
- Continual Learning through Retrieval and ImaginationZhen Wang, Liu Liu, Yiqun Duan, Dacheng TaoAAAI 2022 · 45 citations
- C2MR: Continual Cross-Modal Retrieval for Streaming Multi-modal DataHuaiwen Zhang, Yang Yang, Fan Qi, Shengsheng Qian et al.ACM MM 2023 · 8 citations
- Lifelong GAN: Continual Learning for Conditional Image GenerationMengyao Zhai, Lei Chen, Frederick Tung, Jiawei He et al.ICCV 2019 · 204 citations
- SDDGR: Stable Diffusion-Based Deep Generative Replay for Class Incremental Object DetectionJunsu Kim, Hoseong Cho, Jihyeon Kim, Yihalem Yimolal Tiruneh et al.CVPR 2024
- RATT: Recurrent Attention to Transient Tasks for Continual Image CaptioningRiccardo Del Chiaro, Bartlomiej Twardowski, Andrew D. Bagdanov, Joost van de WeijerNeurIPS 2020 · 55 citations
