Test-time Adaptation for Cross-modal Retrieval with Query Shift
Haobin Li, Peng Hu, Qianjun Zhang, Xi Peng, Xiting Liu, Mouxing Yang
摘要
The success of most existing cross-modal retrieval methods heavily relies on the assumption that the given queries follow the same distribution of the source domain. However, such an assumption is easily violated in real-world scenarios due to the complexity and diversity of queries, thus leading to the query shift problem. Specifically, query shift refers to the online query stream originating from the domain that follows a different distribution with the source one. In this paper, we observe that query shift would not only diminish the uniformity (namely, withinmodality scatter) of the query modality but also amplify the gap between query and gallery modalities. Based on the observations, we propose a novel method dubbed Test-time adaptation for Cross-modal Retrieval (TCR). In brief, TCR employs a novel module to refine the query predictions (namely, retrieval results of the query) and a joint objective to prevent query shift from disturbing the common space, thus achieving online adaptation for the cross-modal retrieval models with query shift. Expensive experiments demonstrate the effectiveness of the proposed TCR against query shift. Code is available at https://github.com/XLearning-SCU/2025-ICLR-TCR . * Corresponding author. (a)Dominant Paradigm (b)Query Shift (c)Observations Pre-trained Models Fine-tune Zero-shot Rib knit wool beanie in 'anthracite' grey. Rubberized logo in red and white at side. Query from Personalized Domains Query from Scarce Domains Long sleeve cotton twill blazer in off-white. Raw edges throughout. Notched lapel collar.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Interactive Cross-modal Learning for Text-3D Scene RetrievalYanglin Feng, Yongxiang Li, Yuan Sun, Yang Qin 等NeurIPS 2025 · 被引用 9 次
- LargeMvC-Net: Anchor-based Deep Unfolding Network for Large-scale Multi-view ClusteringShide Du, Chunming Wu, Zihan Fang, Wendi Zhao 等ACM MM 2025 · 被引用 7 次
- MoRA: Missing Modality Low-Rank Adaptation for Visual RecognitionShu Zhao, Nilesh A. Ahuja, Tan Yu, Tianyi Shen 等ICLR 2026 · 被引用 5 次
- Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue ConsistencyZhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu 等CVPR 2026 · 被引用 3 次
- Diffusion-calibrated Continual Test-time AdaptationXu Yang, Moqi Li, Kun WeiAAAI 2026
它引用的顶会 Paper34
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Test-time Adaptation against Multi-modal Reliability BiasMouxing Yang, Yunfan Li, Changqing Zhang, Peng Hu 等ICLR 2024 · 被引用 41 次
- Test-Time Selective Adaptation for Uni-Modal Distribution Shift in Multi-Modal DataMingcai Chen, Baoming Zhang, Zongbo Han, Wenyu Jiang 等ICML 2025
- Bridging Modalities via Progressive Re-alignment for Multimodal Test-Time AdaptationJiacheng Li, Songhe FengAAAI 2026 · 被引用 2 次
- Attention Bootstrapping for Multi-Modal Test-Time AdaptationYusheng Zhao, Junyu Luo, Xiao Luo, Jinsheng Huang 等AAAI 2025 · 被引用 5 次
- Adapting Point Cloud Analysis via Multimodal Bayesian Distribution LearningXingyu Zhu, Yi Liang, Shuo Wang, Wenbo Zhu 等CVPR 2026
