Universal Weighting Metric Learning for Cross-Modal Matching
Jiwei Wei, Xing Xu, Yang Yang, Yanli Ji, Zheng Wang, Heng Tao Shen
摘要
Cross-modal matching has been a highlighted research topic in both vision and language areas. Learning appropriate mining strategy to sample and weight informative pairs is crucial for the cross-modal matching performance. However, most existing metric learning methods are developed for unimodal matching, which is unsuitable for cross-modal matching on multimodal data with heterogeneous features. To address this problem, we propose a simple and interpretable universal weighting framework for cross-modal matching, which provides a tool to analyze the interpretability of various loss functions. Furthermore, we introduce a new polynomial loss under the universal weighting framework, which defines a weight function for the positive and negative informative pairs respectively. Experimental results on two image-text matching benchmarks and two video-text matching benchmarks validate the efficacy of the proposed method. The source code is available at: https://github.com/wayne980/PolyLoss .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Channel Augmented Joint Learning for Visible-Infrared RecognitionMang Ye, Weijian Ruan, Bo Du, Mike Zheng ShouICCV 2021 · 被引用 310 次
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 被引用 185 次
- HANet: Hierarchical Alignment Networks for Video-Text RetrievalPeng Wu, Xiangteng He, Mingqian Tang, Yiliang Lv 等ACM MM 2021 · 被引用 62 次
- Telling the What while Pointing to the Where: Multimodal Queries for Image RetrievalSoravit Changpinyo, Jordi Pont-Tuset, Vittorio Ferrari, Radu SoricutICCV 2021 · 被引用 30 次
- One-shot Scene Graph GenerationYuyu Guo, Jingkuan Song, Lianli Gao, Heng Tao ShenACM MM 2020 · 被引用 26 次
它引用的顶会 Paper1
相关 Paper
- Meta Self-Paced Learning for Cross-Modal MatchingJiwei Wei, Xing Xu, Zheng Wang, Guoqing WangACM MM 2021 · 被引用 34 次
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
- COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-training for Vision-Language RepresentationKeyu Wen, Jin Xia, Yuanyuan Huang, Linyang Li 等ICCV 2021 · 被引用 35 次
- SOLAR: Self-supervised Joint Learning for Symmetric Multimodal RetrievalWenjie Yang, Hang Yu, Yuyu Guo, Peng DiICML 2026
- Contrastive Learning with Complex HeterogeneityLecheng Zheng, Jinjun Xiong, Yada Zhu, Jingrui HeKDD 2022 · 被引用 29 次
