Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal Reasoning
Yang Liu, Wentao Feng, Shudong Huang, Yalan Ye, Jiancheng Lv
Abstract
Large-scale web-harvested datasets have fueled the progress of cross-modal retrieval but inevitably suffer from noisy correspondence, which severely degrades model generalization. Existing methods primarily address this by filtering out noise or seeking a substitute label, yet they predominantly remain bound by a "Discrete Selection" paradigm. We argue that relying on a single discrete proxy induces Single-Point Fragility and Discretization Error. To overcome these limitations, we propose a novel framework, Intra-modal Neighbor-aware Noise Rectification (IN 2 R), which shifts the paradigm from searching for a substitute to synthesizing a reliable supervision target. Leveraging the intrinsic geometric stability of intra-modal data, IN 2 R employs a Graph Refiner to perform relational reasoning over neighbors retrieved from a dynamic Cross-Model Memory. Instead of propagating discrete labels, our method synthesizes a continuous, soft prototype that reflects the consensus of the local semantic neighborhood, effectively rectifying inter-modal misalignment. Extensive experiments on Flickr30K, MS-COCO, and CC152K demonstrate that IN 2 R significantly outperforms state-of-the-art methods. Our code and pretrained models are publicly available at https: //github.com/liuyyy111/IN2R .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d4ea5f5-ba38-4728-a42e-2daefb6b0c2dBuilds on22
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- CAMP: Cross-Modal Adaptive Message Passing for Text-Image RetrievalZihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng et al.ICCV 2019 · 349 citations
- Learning with Noisy Correspondence for Cross-modal MatchingZhenyu Huang, Guocheng Niu, Xiao Liu, Wenbiao Ding et al.NeurIPS 2021 · 215 citations
- Dynamic Modality Interaction Modeling for Image-Text RetrievalLeigang Qu, Meng Liu, Jianlong Wu, Zan Gao et al.SIGIR 2021 · 187 citations
Related papers
- Noise Self-Correction via Relation Propagation for Robust Cross-Modal RetrievalRuoxuan Li, Xiangyu Wu, Yang YangACM MM 2025
- Noise-Robust Cross-modal Learning for Reliable 2D-3D RetrievalAo Yang, Yanglin Feng, Yuan Sun, Dezhong Peng et al.ACM MM 2025
- PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence LearningZhuoyao Liu, Yang Liu, Wentao Feng, Shudong HuangAAAI 2026
- Learnable Pillar-based Re-ranking for Image-Text RetrievalLeigang Qu, Meng Liu, Wenjie Wang, Zhedong Zheng et al.SIGIR 2023 · 23 citations
- ReCon: Enhancing True Correspondence Discrimination through Relation Consistency for Robust Noisy Correspondence LearningQuanxing Zha, Xin Liu, Shu-Juan Peng, Yiu-ming Cheung et al.CVPR 2025
