Multi-Level Confidence Learning for Trustworthy Multimodal Classification
Xiao Zheng, Chang Tang, Zhiguo Wan, Chengyu Hu, Wei Zhang
Abstract
With the rapid development of various data acquisition technologies, more and more multimodal data come into being. It is important to integrate different modalities which are with high-dimensional features for boosting final multimodal data classification task. However, existing multimodal classification methods mainly focus on exploiting the complementary information of different modalities, while ignoring the learning confidence during information fusion. In this paper, we propose a trustworthy multimodal classification network via multi-level confidence learning, referred to as MLCLNet. Considering that a large number of feature dimensions could not contribute to final classification performance but disturb the discriminability of different samples, we propose a feature confidence learning mechanism to suppress some redundant features, as well as enhancing the expression of discriminative feature dimensions in each modality. In order to capture the inherent sample structure information implied in each modality, we design a graph convolutional network branch to learn the corresponding structure preserved feature representation and generate modal-specific initial classification labels. Since samples from different modalities should share consistent labels, a cross-modal label fusion module is deployed to capture the label correlations of different modalities. In addition, motivated the ideally orthogonality of final fused label matrix, we design a label confidence loss to supervise the network for learning more separable data representations. To the best of our knowledge, MLCLNet is the first work which integrates both feature and label-level confidence learning for multimodal classification. Extensive experiments on four multimodal medical datasets are conducted to validate superior performance of MLCLNet when compared to other state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fa24262-c349-41d2-9fd3-14f744b3ac02Cited by top-tier papers7
- Facilitating Multimodal Classification via Dynamically Learning Modality GapYang Yang, Fengqiang Wan, Qing-Yuan Jiang, Yi XuNeurIPS 2024 · 65 citations
- Self-supervised Trusted Contrastive Multi-view Clustering with Uncertainty RefinedShizhe Hu, Binyan Tian, Weibo Liu, Yangdong YeAAAI 2025 · 11 citations
- Fusion Meets Diverse Conditions: A High-Diversity Benchmark and Baseline for UAV-Based Multimodal Object Detection with Condition CuesChen Chen, Kangcheng Bin, Ting Hu, Jiahao Qi et al.ICCV 2025 · 8 citations
- Trusted Multi-view Learning for Long-tailed ClassificationChuanqing Tang, Yifei Shi, Guanghao Lin, Lei Xing et al.AAAI 2026 · 1 citation
- Who Should I Trust? Explicit Confidence-Focused Multimodal Intent RecognitionYi Liu, Qimeng Yang, Lanlan LuAAAI 2026
Builds on10
- What Makes Multi-Modal Learning Better than Single (Provably)Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen et al.NeurIPS 2021 · 404 citations
- SMIL: Multimodal Learning with Severely Missing ModalityMengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov et al.AAAI 2021 · 393 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
- Multimodal Dynamics: Dynamical Fusion for Trustworthy Multimodal ClassificationZongbo Han, Fan Yang, Junzhou Huang, Changqing Zhang et al.CVPR 2022 · 149 citations
- Generative Multi-View Human Action RecognitionLichen Wang, Zhengming Ding, Zhiqiang Tao, Yunyu Liu et al.ICCV 2019 · 112 citations
Related papers
- DPNET: Dynamic Poly-attention Network for Trustworthy Multi-modal ClassificationXin Zou, Chang Tang, Xiao Zheng, Zhenglai Li et al.ACM MM 2023 · 16 citations
- Heterogeneous Graph Learning for Multi-Modal Medical Data AnalysisSein Kim, Namkyeong Lee, Junseok Lee, Dongmin Hyun et al.AAAI 2023 · 54 citations
- Calibrating Multimodal LearningHuan Ma, Qingyang Zhang, Changqing Zhang, Bingzhe Wu et al.ICML 2023 · 42 citations
- Learning a Graph Neural Network with Cross Modality Interaction for Image FusionJiawei Li, Jiansheng Chen, Jinyuan Liu, Huimin MaACM MM 2023 · 85 citations
- Modality to Modality Translation: An Adversarial Representation Learning and Graph Fusion Network for Multimodal FusionSijie Mai, Haifeng Hu, Songlong XingAAAI 2020 · 233 citations
