Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery Banks
Yimu Wang, Xiangru Jian, Bo Xue
摘要
In this work, we present a post-processing solution to address the hubness problem in cross-modal retrieval, a phenomenon where a small number of gallery data points are frequently retrieved, resulting in a decline in retrieval performance. We first theoretically demonstrate the necessity of incorporating both the gallery and query data for addressing hubness as hubs always exhibit high similarity with gallery and query data. Second, building on our theoretical results, we propose a novel framework, Dual Bank Normalization (DBNORM). While previous work has attempted to alleviate hubness by only utilizing the query samples, DBNORM leverages two banks constructed from the query and gallery samples to reduce the occurrence of hubs during inference. Next, to complement DBNORM, we introduce two novel methods, dual inverted softmax and dual dynamic inverted softmax, for normalizing similarity based on the two banks. Specifically, our proposed methods reduce the similarity between hubs and queries while improving the similarity between non-hubs and queries. Finally, we present extensive experimental results on diverse language-grounded benchmarks, including text-image, text-video, and text-audio, demonstrating the superior performance of our approaches compared to previous methods in addressing hubness and boosting retrieval performance. Our code is available at https://github.com/yimuwangcs/Better_Cross_Modal_Retrieval. ©2023 Association for Computational Linguistics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Fewer Steps, Better Performance: Efficient Cross-Modal Clip Trimming for Video Moment Retrieval Using LanguageXiang Fang, Daizong Liu, Wanlong Fang, Pan Zhou 等AAAI 2024 · 被引用 30 次
- LLM-Enhanced Action-Aware Multi-Modal Prompt Tuning for Image-Text MatchingMengxiao Tian, Xinxiao Wu, Shuo YangICCV 2025 · 被引用 3 次
- Rebalancing Contrastive Alignment with Bottlenecked Semantic Increments in Text-Video RetrievalJian Xiao, Zijie Song, Jialong Hu, Hao Cheng 等NeurIPS 2025 · 被引用 3 次
- Hubness Reduction with Dual Bank Sinkhorn Normalization for Cross-Modal RetrievalZhengxin Pan, Haishuai Wang, Fangyu Wu, Peng Zhang 等ACM MM 2025 · 被引用 2 次
- Prediction Hubs are Context-Informed Frequent Tokens in LLMsBeatrix Miranda Ginn Nielsen, Iuri Macocco, Marco BaroniACL 2025 · 被引用 2 次
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
相关 Paper
- Cross Modal Retrieval with Querybank NormalisationSimion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu 等CVPR 2022 · 被引用 84 次
- NeighborRetr: Balancing Hub Centrality in Cross-Modal RetrievalZengrong Lin, Zheng Wang, Tianwen Qian, Pan Mu 等CVPR 2025
- One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via HubnessHiroyuki Deguchi, Katsuki Chousa, Yusuke SakaiACL 2026
- HAL: Improved Text-Image Matching by Mitigating Visual Semantic HubsFangyu Liu, Rongtian Ye, Xun Wang, Shuaipeng LiAAAI 2020 · 被引用 36 次
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video RetrievalDohwan Ko, Ji Soo Lee, Minhyuk Choi, Zihang Meng 等ICCV 2025 · 被引用 4 次
