Universal Weighting Metric Learning for Cross-Modal Matching
Jiwei Wei, Xing Xu, Yang Yang, Yanli Ji, Zheng Wang, Heng Tao Shen
Abstract
Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most existing metric learning methods are developed for unimodal matching, which is unsuitable for cross-modal matching on multimodal data with heterogeneous features. To address this problem, we propose a simple and interpretable universal weighting framework for cross-modal matching, which provides a tool to analyze the interpretability of various loss functions. Furthermore, we introduce a new polynomial loss under the universal weighting framework, which defines a weight function for the positive and negative informative pairs respectively. Experimental results on two image-text matching benchmarks and two video-text matching benchmarks validate the efficacy of the proposed method. The source code is available at: https://github.com/wayne980/PolyLoss .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Channel Augmented Joint Learning for Visible-Infrared RecognitionMang Ye, Weijian Ruan, Bo Du, Mike Zheng ShouICCV 2021 · 310 citations
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 185 citations
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv et al.ACM MM 2021 · 62 citations
- Telling the What while Pointing to the Where: Multimodal Queries for Image RetrievalSoravit Changpinyo, Jordi Pont-Tuset, Vittorio Ferrari, Radu SoricutICCV 2021 · 30 citations
- One-shot Scene Graph GenerationYuyu Guo, Jingkuan Song, Lianli Gao, Heng Tao ShenACM MM 2020 · 26 citations
Builds on1
Related papers
- Meta Self-Paced Learning for Cross-Modal MatchingJiwei Wei, Xing Xu, Zheng Wang, Guoqing WangACM MM 2021 · 34 citations
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
- COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language RepresentationKeyu Wen, Jin Xia, Yuanyuan Huang, Linyang Li et al.ICCV 2021 · 35 citations
- SOLAR: Self-supervised Joint Learning for Symmetric Multimodal RetrievalWenjie Yang, Hang Yu, Yuyu Guo, Peng DiICML 2026
- Contrastive Learning with Complex HeterogeneityLecheng Zheng, Jinjun Xiong, Yada Zhu, Jingrui HeKDD 2022 · 29 citations
