C2MR: Continual Cross-Modal Retrieval for Streaming Multi-modal Data
Huaiwen Zhang, Yang Yang, Fan Qi, Shengsheng Qian, Changsheng Xu
Abstract
Massive numbers of new images are uploaded to the internet every day. However, existing cross-modal retrieval (CMR) approaches struggle to accommodate this continuously growing data. The prevalent practice involves periodically retraining or fine-tuning a new model based on the accumulated data, which in turn invalidates billions of indexed features extracted by the previous model and incurs another substantial computational cost to extract new features for the entire data archive. Is it possible to develop a retrieval model that effectively captures the knowledge of upcoming sessions while preserving the discriminative power of features extracted in previous sessions? In this paper, we propose an online continual learning setup, OC-CMR, to formalize the data-incremental growth challenge faced by cross-modal retrieval systems. It consists of two key settings: 1) Similar to the real-world scenarios, the streaming multi-modal data arrives once per session; 2) Consider the computational costs, each instance of archived data has its feature extracted only once and by its corresponding model in its session. Based on our OC-CMR, we perform in-depth evaluations of state-of-the-art cross-modal retrieval methods and observe that they suffer from representational shift and collapse due to the catastrophic forgetting. To address this issue, we propose the Continual Cross-Modal Retrieval (C2MR) approach, which learns a shared common space not only across modalities but also sessions and maintains relationships between samples from distinct sessions via cross-modal relational coherence and semantic representation coordination. We construct two new benchmarks by adapting MS-COCO and Flickr30K datasets to the OC-CMR setting, providing a more challenging evaluation framework for CMR tasks. Experimental results demonstrate that our method effectively alleviates forgetting and significantly outperforms combinations of previous arts in cross-modal retrieval and continual learning.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get dad37426-eb4b-430c-a361-684b9fb898aaRelated papers
- Online Cross-Modal Hashing with Expanding Label SpaceWentao Fan, Chao Zhang, Chunlin Chen, Huaxiong LiAAAI 2026
- Knowledge Decomposition and Replay: A Novel Cross-modal Image-Text Retrieval Continual Learning MethodRui Yang, Shuang Wang, Huan Zhang, Siyuan Xu et al.ACM MM 2023 · 13 citations
- BMU-MoCo: Bidirectional Momentum Update for Continual Video-Language ModelingYizhao Gao, Nanyi Fei, Haoyu Lu, Zhiwu Lu et al.NeurIPS 2022 · 4 citations
- Towards Multimodal Continual Knowledge Embedding wth Modality Forgetting ModulationXiaowen Jiang, Jing Yang, Shundong Yang, Yuan Gao et al.AAAI 2026
- Online Cross-Modal Hashing with Multi-Level MemoryWentao Fan, Chao Zhang, Chunlin Chen, Huaxiong LiACM MM 2025
