TriSim: Tri-Dimensional Similarity Modeling with Extreme Value Theory for False-Negative Mitigation in Remote Sensing Image-Text Retrieval
Chengyu Zheng, Hanzhang Lu, Jie Nie, Shan Du
摘要
In remote sensing (RS) cross-modal retrieval, most existing methods employ contrastive learning as their primary optimization objective, aligning anchors with positive counterparts and distinguishing them from negative samples. To improve negative sampling, these approaches typically set thresholds on cross-modal similarity scores, designating negatives that exceed the threshold as false negative samples (FNS). However, dependence on a single cross-modal similarity threshold is fragile because it fails to account for the cross-modal semantic overlaps and gaps. To address these challenges, we introduce TriSim, a novel image-text retrieval framework that constructs a tri-dimensional negative similarity space <img-img, img-txt, txt-txt> to mitigate the influence of the FNS issue. Specifically, considering that FNS appear as anomalies in this space, two complementary strategies are developed to characterize the statistical properties of the tail distribution for FNS selection: one identifies samples distant from the dense ellipsoidal center, and the other targets upper-right high-similarity extremes. The FNS from both strategies are unified, followed by Bernoulli sampling to guide the triplet loss optimization. To further refine the selected FNS, intra-modal saliency differences are computed to generate masks that guide the learning of a gain matrix, which amplifies highly discriminative regions and suppresses ambiguous ones. Extensive experiments on two benchmarks demonstrate the superiority of the proposed TriSim in mitigating the influence of false negatives in RS image-text retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
- MedCLIP: Contrastive Learning from Unpaired Medical Images and TextZifeng Wang, Zhenbang Wu, Dinesh Agarwal, Jimeng SunEMNLP 2022 · 被引用 907 次
- Similarity Reasoning and Filtration for Image-Text MatchingHaiwen Diao, Ying Zhang, Lin Ma, Huchuan LuAAAI 2021 · 被引用 413 次
相关 Paper
- Your Negative May not Be True Negative: Boosting Image-Text Matching with False Negative EliminationHaoxuan Li, Yi Bin, Junrong Liao, Yang Yang 等ACM MM 2023 · 被引用 42 次
- Robust Remote Sensing Image–Text Retrieval with Noisy Correspondenceqiya song, Yiqiang Xie, Yuan Sun, Renwei Dian 等CVPR 2026 · 被引用 1 次
- PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image-Text RetrievalPengxiang Ouyang, Qing Ma, Zheng Wang, Cong BaiAAAI 2026
- Entity-Level Alignment with Prompt-Guided Adapter for Remote Sensing Image-Text RetrievalShuoshuo Li, Shuli Cheng, Liejun WangACM MM 2025 · 被引用 2 次
- TriSampler: A Better Negative Sampling Principle for Dense RetrievalZhen Yang, Zhou Shao, Yuxiao Dong, Jie TangAAAI 2024 · 被引用 17 次
