Fuzzy Multimodal Learning for Trusted Cross-modal Retrieval
Siyuan Duan, Yuan Sun, Dezhong Peng, Zheng Liu, Xiaomin Song, Peng Hu
Abstract
Cross-modal retrieval aims to match related samples across distinct modalities, facilitating the retrieval and discovery of heterogeneous information. Although existing methods show promising performance, most are deterministic models and are unable to capture the uncertainty inherent in the retrieval outputs, leading to potentially unreliable results. To address this issue, we propose a novel framework called FUzzy Multimodal lEarning (FUME), which is able to self-estimate epistemic uncertainty, thereby embracing trusted cross-modal retrieval. Specifically, our FUME leverages the Fuzzy Set Theory to view the outputs of the classification network as a set of membership degrees and quantify category credibility by incorporating both possibility and necessity measures. However, directly optimizing the category credibility could mislead the model by overoptimizing the necessity for unmatched categories. To overcome this challenge, we present a novel fuzzy multimodal learning strategy, which utilizes label information to guide necessity optimization in the right direction, thereby indirectly optimizing category credibility and achieving accurate decision uncertainty quantification. Furthermore, we design an uncertainty merging scheme that accounts for decision uncertainties, thus further refining uncertainty estimates and boosting the trustworthiness of retrieval results. Extensive experiments on five benchmark datasets demonstrate that FUME remarkably improves both retrieval performance and reliability, offering a prospective solution for cross-modal retrieval in high-stakes applications. Code is available at https://github.com/siyuancncd/FUME .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dfac9f86-4642-414c-99c8-f35e216e052dCited by top-tier papers13
- Robust Cross-modal Alignment Learning for Cross-Scene Spatial Reasoning and GroundingYanglin Feng, Hongyuan Zhu, Dezhong Peng, Xi Peng et al.NeurIPS 2025 · 6 citations
- Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video RetrievalJun Li, Xuhang Lou, Jinpeng Wang, Yuting Wang et al.CVPR 2026 · 3 citations
- Deep Unsupervised Hashing via External GuidanceQihong Song, XitingLiu, Hongyuan Zhu, Joey Tianyi Zhou et al.ICML 2025
- Deep Fuzzy Multi-view Learning for Reliable ClassificationSiyuan Duan, Yuan Sun, Dezhong Peng, Guiduo Duan et al.ICML 2025
- Stationary and Clustering Transformer Hashing for Cross-modal RetrievalZhan Yang, Yiran Liu, Youyuan Huang, Yinan LiAAAI 2026
Builds on11
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Deep Joint-Semantics Reconstructing Hashing for Large-Scale Unsupervised Cross-Modal RetrievalShupeng Su, Zhisheng Zhong, Chao ZhangICCV 2019 · 261 citations
- Deep Evidential Learning with Noisy Correspondence for Cross-modal RetrievalYang Qin, Dezhong Peng, Xi Peng, Xu Wang et al.ACM MM 2022 · 101 citations
- Dual Self-Paced Cross-Modal HashingYuan Sun, Jian Dai, Zhenwen Ren, Yingke Chen et al.AAAI 2024 · 35 citations
- Semi-Supervised Multi-Modal Learning with Balanced Spectral DecompositionPeng Hu, Hongyuan Zhu, Xi Peng, Jie LinAAAI 2020 · 28 citations
Related papers
- Deep Evidential Hashing for Trustworthy Cross-Modal RetrievalYuan Li, Liangli Zhen, Yuan Sun, Dezhong Peng et al.AAAI 2025 · 8 citations
- Prototype-based Aleatoric Uncertainty Quantification for Cross-modal RetrievalHao Li, Jingkuan Song, Lianli Gao, Xiaosu Zhu et al.NeurIPS 2023 · 52 citations
- Deep Probabilistic Binary Embedding via Learning Reliable Uncertainty for Cross-Modal RetrievalKun Cheng, Qibing Qin, Wenfeng Zhang, Lei Huang et al.ACM MM 2025 · 4 citations
- FUSE: Quantifying Uncertainty in Vision-Language Models by Bayesian Fusing Epistemic and Aleatoric UncertaintyHarry Zhang, Luca CarloneICML 2026
- Probabilistic Embeddings for Cross-Modal RetrievalSanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis et al.CVPR 2021
