Robust Noisy Correspondence Learning with Equivariant Similarity Consistency
Yuchen Yang, Likai Wang, Erkun Yang, Cheng Deng
摘要
The surge in multimodal data has propelled cross-modal matching to the forefront of research interest. However, the challenge lies in the laborious and expensive process of curating a large and accurately matched multimodal dataset. Commonly sourced from the Internet, these datasets often suffer from a significant presence of mismatched data, impairing the performance of matching models. To address this problem, we introduce a novel regularization approach named Equivariant Similarity Con-sistency (ESC), which can facilitate robust clean and noisy data separation and improve the training for cross-modal matching. Intuitively, our method posits that the semantic variations caused by image changes should be proportional to those caused by text changes for any two matched samples. Accordingly, we first calculate the ESC by comparing image and text semantic variations between a set of elab-orated anchor points and other undivided training data. Then, pairs with high ESC are filtered out as noisy correspondence pairs. We implement our method by combining the ESC with a traditional hinge-based triplet loss. Exten-sive experiments on three widely used datasets, including Flickr30K, MS-COCO, and Conceptual Captions, verify the effectiveness of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Meta-Learning Dynamic Center Distance: Hard Sample Mining for Learning with Noisy LabelsChenyu Mu, Yijun Qu, Jiexi Yan, Erkun Yang 等ICCV 2025 · 被引用 2 次
- TSVC: Tripartite Learning with Semantic Variation Consistency for Robust Image-Text RetrievalShuai Lyu, Zijing Tian, Zhonghong Ou, Yifan Zhu 等AAAI 2025 · 被引用 2 次
- ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence LearningQuanxing Zha, Xin Liu, Shu-Juan Peng, Yiu-ming Cheung 等CVPR 2025
- PAUL: Uncertainty-Guided Partition and Augmentation for Robust Cross-View Geo-Localization under Noisy CorrespondenceZheng Li, Xueyi Zhang, Yanming Guo, Yuxiang Xie 等CVPR 2026
- Noise Self-Correction via Relation Propagation for Robust Cross-Modal RetrievalRuoxuan Li, Xiangyu Wu, Yang YangACM MM 2025
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Simple and Deep Graph Convolutional NetworksMing Chen, Zhewei Wei, Zengfeng Huang, Bolin Ding 等ICML 2020 · 被引用 1,910 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li 等ICCV 2019 · 被引用 598 次
相关 Paper
- Noisy Correspondence Rectification via Asymmetric Similarity LearningYunbo Wang, YuJie Wu, Zhien Dai, Can Tian 等AAAI 2025 · 被引用 4 次
- Deep Evidential Learning with Noisy Correspondence for Cross-modal RetrievalYang Qin, Dezhong Peng, Xi Peng, Xu Wang 等ACM MM 2022 · 被引用 101 次
- Deep Evidential Hashing for Trustworthy Cross-Modal RetrievalYuan Li, Liangli Zhen, Yuan Sun, Dezhong Peng 等AAAI 2025 · 被引用 8 次
- Noisy Correspondence Learning with Meta Similarity CorrectionHaochen Han, Kaiyao Miao, Qinghua Zheng, Minnan LuoCVPR 2023
- Multimodal Aligned Semantic Knowledge for Unpaired Image-text MatchingLaiguo Yin, Yixin Zhang, YUQING SUN, Lizhen CuiICLR 2026
