Meta Self-Paced Learning for Cross-Modal Matching
Jiwei Wei, Xing Xu, Zheng Wang, Guoqing Wang
Abstract
Cross-modal matching has attracted growing attention due to the rapid emergence of the multimedia data on the web and social applications. Recently, many re-weighting methods have been proposed for accelerating model training by designing a mapping function from similarity scores to weights. However, these re-weighting methods are difficult to be universally applied in practice since manually pre-set weighting functions inevitably involve hyper-parameters. In this paper, we propose a Meta Self-Paced Network (Meta-SPN) that automatically learns a weighting scheme from data for cross-modal matching. Specifically, a meta self-paced network composed of a fully connected neural network is designed to fit the weight function, which takes the similarity score of the sample pairs as input and outputs the corresponding weight value. Our meta self-paced network considers not only the self-similarity scores, but also their potential interactions (e.g., relative-similarity) when learning the weights. Motivated by the success of meta-learning, we use the validation set to update the meta self-paced network during the training of the matching network. Experiments on two image-text matching benchmarks and two video-text matching benchmarks demonstrate the generalization and effectiveness of our method.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers6
- Cross-Lingual Cross-Modal Retrieval with Noise-Robust LearningYabing Wang, Jianfeng Dong, Tianxiang Liang, Minsong Zhang et al.ACM MM 2022 · 26 citations
- Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsRuitao Pu, Yuan Sun, Yang Qin, Zhenwen Ren et al.AAAI 2025 · 25 citations
- KDA: Knowledge Diffusion Alignment with Enhanced Context for Video Temporal GroundingRan Ran, Jiwei Wei, Shiyuan He, Zeyu Ma et al.ICCV 2025 · 4 citations
- Open-Scenario Domain Adaptive Object Detection in Autonomous DrivingZeyu Ma, Ziqiang Zheng, Jiwei Wei, Xiaoyong Wei et al.ACM MM 2023 · 2 citations
- Robust Remote Sensing Image–Text Retrieval with Noisy Correspondenceqiya song, Yiqiang Xie, Yuan Sun, Renwei Dian et al.CVPR 2026 · 1 citation
Related papers
- Universal Weighting Metric Learning for Cross-Modal MatchingJiwei Wei, Xing Xu, Yang Yang, Yanli Ji et al.CVPR 2020
- Adaptable Text Matching via Meta-Weight RegulatorBo Zhang, Chen Zhang, Fang Ma, Dawei SongSIGIR 2022
- Noisy Correspondence Learning with Meta Similarity CorrectionHaochen Han, Kaiyao Miao, Qinghua Zheng, Minnan LuoCVPR 2023
- Meta-Guided Sample Reweighting for Robust Cross-Modal Hashing Retrieval with Noisy LabelsZiang Tan, Weitao An, Erkun YangAAAI 2026
- Meta-Guided Adaptive Weight Learner for Noisy CorrespondenceChenyu Mu, Erkun Yang, Cheng DengSIGIR 2025 · 2 citations
