Hyper-Opinion Vagueness Quantification for Robust Multimodal Learning
Disen Hu, Xun Jiang, Xiaofeng Cao, Zheng Wang, Jingkuan Song, Heng Tao Shen, Xing Xu
Abstract
Robust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``uncertain'' in extracting discriminative features of a multimodal model as vagueness. Neglecting such vagueness, as previous RML works commonly do, will undermine the ability to extract unique semantics of each category in multimodal models, further resulting in worse robustness under disturbances that affect semantic representations. Additionally, this vagueness will lead the parameter updating processes towards unreliable fusion, thus diverting the learning processes of the multimodal model from learning unique features of each category. Based on the above insight, we propose a novel robust multimodal learning approach, termed Hyper-Opinion Quantifying Vagueness (HOQV). Specifically, we first introduce hyper-opinion to capture and quantify the vagueness of multimodal learning in discriminating representations of different categories. Moreover, to mitigate the interference in parameter updating of unreliable representations with high vagueness, we also design the Hyper-Opinion Gradient Modulation to guide the optimization processes. We evaluate our HOQV on six datasets with different disturbances, including noise and adversarial attack, and demonstrate that our proposed method achieves state-of-the-art performance consistently.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc33e7ea-8bfa-4f95-a58f-c2fb53151b90Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Evidential Deep Learning for Open Set Action RecognitionWentao Bao, Qi Yu, Yu KongICCV 2021 · 204 citations
- Uncertainty Aware Semi-Supervised Learning on Graph DataXujiang Zhao, Feng Chen, Shu Hu, Jin-Hee ChoNeurIPS 2020 · 178 citations
- ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment AnalysisJiuding Yang, Yakun Yu, Di Niu, Weidong Guo et al.ACL 2023 · 135 citations
- Reliable Conflictive Multi-View LearningCai Xu, Jiajun Si, Ziyu Guan, Wei Zhao et al.AAAI 2024 · 121 citations
- MMPareto: Boosting Multimodal Learning with Innocent Unimodal AssistanceYake Wei, Di HuICML 2024 · 86 citations
Related papers
- Hyper-opinion Evidential Deep Learning for Out-of-Distribution DetectionJingen Qu, Yufei Chen, Xiaodong Yue, Wei Fu et al.NeurIPS 2024 · 17 citations
- Balanced Multimodal Learning via On-the-fly Gradient ModulationXiaokang Peng, Yake Wei, Andong Deng, Dong Wang et al.CVPR 2022 · 264 citations
- Hyper Evidential Deep Learning to Quantify Composite Classification UncertaintyChangbin Li, Kangshuo Li, Yuzhe Ou, Lance M. Kaplan et al.ICLR 2024 · 10 citations
- Dynamic Evidence Decoupling for Trusted Multi-view LearningYing Liu, Lihong Liu, Cai Xu, Xiangyu Song et al.ACM MM 2024 · 11 citations
- Vulnerability-Aware Robust Multimodal Adversarial TrainingJunrui Zhang, Xinyu Zhao, Jie Peng, Chenjie Wang et al.AAAI 2026
