Noise-Robust Cross-modal Learning for Reliable 2D-3D Retrieval
Ao Yang, Yanglin Feng, Yuan Sun, Dezhong Peng, Guiduo Duan, Yang Qin
Abstract
With the rapid proliferation of 2D and 3D data, driven by advances in virtual environments and AI-generated content, cross-modal 2D-3D retrieval has attracted growing attention. However, it is easy to introduce noisy labels due to the spatial complexity of 3D content. Although various methods have been proposed to address this issue, they still struggle to handle or effectively re-exploit noisy samples. Moreover, existing approaches are prone to error accumulation due to the self-reinforcement of the model during training. To address these issues, we propose a Noise-Robust Cross-modal Learning (NRCL) framework based on the hybrid strategy. Specifically, NRCL introduces a Robust Cross-modal Co-separator (RCC), which separates noisy samples from clean ones by leveraging modality complementarity and adopting a co-teaching paradigm to mitigate potential error accumulation of the single model during training. Besides, a Reliable Soft Rectification (RSR) method is adopted to correct noisy labels by aggregating historical and dual-model predictions, exploiting the discriminative information from noisy samples. Finally, a Robust Cross-modal Prototype Learning (RCPL) is proposed to improve the discriminability of inter-class and alleviate the inherent gaps across modalities in the shared common space, which jointly leverages clean and rectified labels, thereby mitigating the detrimental impact of noisy samples. Extensive experiments are conducted on three 3D multimodal datasets to verify the effectiveness of our method by comparing it with 10 state-of-the-art methods. The code is available at https://github.com/yangaonidaye123/NRCL.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 166913bb-dc0c-4b1a-a635-0ce4b81d796eCited by top-tier papers4
- Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute MiningHao Zou, Runqing Zhang, Jin Ding, xue zhou et al.CVPR 2026
- Multimodal Nested Learning for Decoupled and Coordinated OptimizationYanglin Feng, Yang Qin, Dezhong Peng, Rui Wang et al.ICML 2026
- Learning with Admissibility: Robust Fuzzy Hashing for Cross-Modal Retrieval with Noisy LabelsXincheng Sun, Ruitao Pu, Guangsi Shi, Zhenwen Ren et al.ICML 2026
- Robust Semi-paired Multimodal Learning for Cross-modal RetrievalYang Qin, Yuan Sun, Xi Peng, Dezhong Peng et al.AAAI 2026
Related papers
- RONO: Robust Discriminative Learning with Noisy Labels for 2D-3D Cross-Modal RetrievalYanglin Feng, Hongyuan Zhu, Dezhong Peng, Xi Peng et al.CVPR 2023
- DAC: 2D-3D Retrieval with Noisy Labels via Divide-and-Conquer Alignment and CorrectionChaofan Gan, Yuanpeng Tu, Yuxi Li, Weiyao LinACM MM 2024 · 2 citations
- Neighbor-aware Instance Refining with Noisy Labels for Cross-Modal RetrievalYizhi Liu, Ruitao Pu, Shilin Xu, Yingke Chen et al.AAAI 2026 · 1 citation
- Robust Contrastive Cross-modal Hashing with Noisy LabelsLongan Wang, Yang Qin, Yuan Sun, Dezhong Peng et al.ACM MM 2024 · 14 citations
- Intra-Modal Neighbors Never Lie: Rectifying Inter-Modal Noisy Correspondence via Graph-Based Intra-Modal ReasoningYang Liu, Wentao Feng, Shudong Huang, Yalan Ye et al.ICML 2026
