Robust Noisy Correspondence Learning with Equivariant Similarity Consistency
Yuchen Yang, Likai Wang, Erkun Yang, Cheng Deng
Abstract
The surge in multimodal data has propelled cross-modal matching to the forefront of research interest. However, the challenge lies in the laborious and expensive process of curating a large and accurately matched multimodal dataset. Commonly sourced from the Internet, these datasets often suffer from a significant presence of mismatched data, impairing the performance of matching models. To address this problem, we introduce a novel regularization approach named Equivariant Similarity Con-sistency (ESC), which can facilitate robust clean and noisy data separation and improve the training for cross-modal matching. Intuitively, our method posits that the semantic variations caused by image changes should be proportional to those caused by text changes for any two matched samples. Accordingly, we first calculate the ESC by comparing image and text semantic variations between a set of elab-orated anchor points and other undivided training data. Then, pairs with high ESC are filtered out as noisy correspondence pairs. We implement our method by combining the ESC with a traditional hinge-based triplet loss. Exten-sive experiments on three widely used datasets, including Flickr30K, MS-COCO, and Conceptual Captions, verify the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10685bf6-55a6-4132-8f9c-1d67f9e7bdc4Cited by top-tier papers11
- Meta-Learning Dynamic Center Distance: Hard Sample Mining for Learning with Noisy LabelsChenyu Mu, Yijun Qu, Jiexi Yan, Erkun Yang et al.ICCV 2025 · 2 citations
- TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text RetrievalShuai Lyu, Zijing Tian, Zhonghong Ou, Yifan Zhu et al.AAAI 2025 · 2 citations
- ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence LearningQuanxing Zha, Xin Liu, Shu-Juan Peng, Yiu-ming Cheung et al.CVPR 2025
- PAUL: Uncertainty-Guided Partition and Augmentation for Robust Cross-View Geo-Localization under Noisy CorrespondenceZheng Li, Xueyi Zhang, Yanming Guo, Yuxiang Xie et al.CVPR 2026
- Noise Self-Correction via Relation Propagation for Robust Cross-Modal RetrievalRuoxuan Li, Xiangyu Wu, Yang YangACM MM 2025
Builds on26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding et al.ICML 2020 · 1,910 citations
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 1,326 citations
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
Related papers
- Noisy Correspondence Rectification via Asymmetric Similarity LearningYunbo Wang, YuJie Wu, Zhien Dai, Can Tian et al.AAAI 2025 · 4 citations
- Deep Evidential Learning with Noisy Correspondence for Cross-modal RetrievalYang Qin, Dezhong Peng, Xi Peng, Xu Wang et al.ACM MM 2022 · 101 citations
- Deep Evidential Hashing for Trustworthy Cross-Modal RetrievalYuan Li, Liangli Zhen, Yuan Sun, Dezhong Peng et al.AAAI 2025 · 8 citations
- Noisy Correspondence Learning with Meta Similarity CorrectionHaochen Han, Kaiyao Miao, Qinghua Zheng, Minnan LuoCVPR 2023
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
