Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided Learning
Zhao Zhou, Weizhong Zhang, Xiangcheng Du, Yingbin Zheng, Cheng Jin
摘要
Image-text matching is a crucial task that bridges visual and linguistic modalities. Recent research typically formulates it into the problem of maximizing the margin with the truly hardest negatives to enhance the learning efficiency and avoid the poor local optima. We argue that such formulation can lead to a serious limitation, i.e., under this formulation, conventional trainers would confine their horizon within the hardest negative examples, while other negative examples offer a range of semantic differences not present in the hardest negatives. In this paper, we propose an efficient negative distribution guided training framework for image-text matching to unlock the substantial promotion space left by the above limitation. Rather than simply incorporating additional negative examples into the training objective, which could diminish both the leading role of the hardest negatives in training and the effect of a large margin learning in producing a robust matching model, our central idea is to supply the objective with distributional information on the entire set of negative examples. To be precise, we first construct the sample similarity matrix based on several pretrained models to extract the distributional information of the entire negative sample dataset. Then we encode it into a margin regularization module to smooth the similarities differences of all negatives. This enhancement facilitates the capture of fine-grained semantic differences and guides the main learning process by maximizing the margin with hard negative examples. Furthermore, we propose a hardest negative rectification module to address the instability in hardest negative selection based on predicted similarity and to correct erroneous hardest negatives. We evaluate our method in combination with several state-of-the-art image-text matching methods, and our quantitative and qualitative experiments demonstrate its significant generalizability and effectiveness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- CompRess: Self-Supervised Learning by Compressing RepresentationsSoroush Abbasi Koohpayegani, Ajinkya Tejankar, Hamed PirsiavashNeurIPS 2020 · 被引用 105 次
- ISD: Self-Supervised Learning by Iterative Similarity DistillationAjinkya Tejankar, Soroush Abbasi Koohpayegani, Vipin Pillai, Paolo Favaro 等ICCV 2021 · 被引用 47 次
相关 Paper
- Synthesizing Counterfactual Samples for Effective Image-Text MatchingHao Wei, Shuhui Wang, Xinzhe Han, Zhe Xue 等ACM MM 2022 · 被引用 11 次
- Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative EliminationHaoxuan Li, Yi Bin, Junrong Liao, Yang Yang 等ACM MM 2023 · 被引用 42 次
- VL-Match: Enhancing Vision-Language Pretraining with Token-Level and Instance-Level MatchingJunyu Bi, Daixuan Cheng, Ping Yao, Bochen Pang 等ICCV 2023 · 被引用 6 次
- Uniformly Distributed Category Prototype-Guided Vision-Language Framework for Long-Tail RecognitionXiaoxuan He, Siming Fu, Xinpeng Ding, Yuchen Cao 等ACM MM 2023 · 被引用 6 次
- Learning Hierarchical Cross-modal Association with Intra-modal Context for Text-Image Person RetrievalYifei Deng, Chenglong Li, Futian Wang, Jin TangACM MM 2025 · 被引用 2 次
