Robust Remote Sensing Image–Text Retrieval with Noisy Correspondence
qiya song, Yiqiang Xie, Yuan Sun, Renwei Dian, Xudong Kang
Abstract
As a pivotal task that bridges remote visual and linguistic understanding, Remote Sensing Image-Text Retrieval (RSITR) has attracted considerable research interest in recent years. However, almost all RSITR methods implicitly assume that image-text pairs are matched perfectly. In practice, acquiring a large set of well-aligned data pairs is often prohibitively expensive or even infeasible. In addition, we also notice that the remote sensing datasets (e.g., RSITMD) truly contain some inaccurate or mismatched image text descriptions. Based on the above observations, we reveal an important but untouched problem in RSITR, i.e., Noisy Correspondence (NC). To overcome these challenges, we propose a novel Robust Remote Sensing Image-Text Retrieval (RRSITR) paradigm that designs a selfpaced learning strategy to mimic human cognitive learning patterns, thereby learning from easy to hard from multimodal data with NC. Specifically, we first divide all training sample pairs into three categories based on the loss magnitude of each pair, i.e., clean sample pairs, ambiguous sample pairs, and noisy sample pairs. Then, we respectively estimate the reliability of each training pair by assigning a weight to each pair based on the values of the loss. Further, we respectively design a new multi-modal self-paced function to dynamically regulate the training sequence and weights of the samples, thus establishing a progressive learning process. Finally, for noisy sample pairs, we present a robust triplet loss to dynamically adjust the soft margin based on semantic similarity, thereby enhancing the robustness against noise. Extensive experiments on three popular benchmark datasets demonstrate that the proposed RRSITR significantly outperforms the state-of-the-art methods, especially in high noise rates. The code is available at: https://github.com/MSFLabX/RRSITR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a239a39-7f12-4ca2-bf84-bb84951d800aBuilds on11
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 413 citations
- Negative-Aware Attention Framework for Image-Text MatchingKun Zhang, Zhendong Mao, Quan Wang, Yongdong ZhangCVPR 2022 · 185 citations
- S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist CaptionsSangwoo Mo, Minkyu Kim, Kyungmin Lee, Jinwoo ShinNeurIPS 2023 · 53 citations
- A Prior Instruction Representation Framework for Remote Sensing Image-text RetrievalJiancheng Pan, Qing Ma, Cong BaiACM MM 2023 · 52 citations
- Dual Self-Paced Cross-Modal HashingYuan Sun, Jian Dai, Zhenwen Ren, Yingke Chen et al.AAAI 2024 · 35 citations
Related papers
- Robust Semi-paired Multimodal Learning for Cross-modal RetrievalYang Qin, Yuan Sun, Xi Peng, Dezhong Peng et al.AAAI 2026
- PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text RetrievalPengxiang Ouyang, Qing Ma, Zheng Wang, Cong BaiAAAI 2026
- Cross-modal Active Complementary Learning with Self-refining CorrespondenceYang Qin, Yuan Sun, Dezhong Peng, Joey Tianyi Zhou et al.NeurIPS 2023 · 49 citations
- HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image RetrievalZixu Li, Yupeng Hu, Zhiwei Chen, Shiqi Zhang et al.AAAI 2026 · 8 citations
- Robust Self-Paced Hashing for Cross-Modal Retrieval with Noisy LabelsRuitao Pu, Yuan Sun, Yang Qin, Zhenwen Ren et al.AAAI 2025 · 25 citations
