Who Should I Trust? Explicit Confidence-Focused Multimodal Intent Recognition
Yi Liu, Qimeng Yang, Lanlan Lu
Abstract
Multimodal intent recognition is aimed at understanding user intentions by integrating information from multiple modalities. It has attracted increasing attention in recently developed dialog systems. The existing studies have focused mainly on modeling semantic interactions within and across modalities, but they often overlook the reliability of each modality. In real-world scenarios, inputs may be corrupted by noisy audio, blurred or occluded videos, or ambiguous text, making it difficult for the employed model to determine who to trust and how much to trust. To address this challenge, we propose a method called explicit confidence-focused multimodal intent recognition (ECFMIR). The core idea of this approach is to assign each modality and each cross-modal associations feature a dedicated confidence lens (CLens) that explicitly estimates the confidence level in a hypothetical manner. This design helps reduce the degree of uncertainty and mitigate the risk of incorrect predictions when addressing conflicting inputs. Comprehensive experiments conducted on two benchmark multimodal intent recognition datasets demonstrate the effectiveness of our method. A further analysis reveals that ECFMIR achieves significant advantages for high-conflict categories and under low-resource conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 079d5ff8-8923-4c6d-a750-2e70c5d04ea6Builds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- MIntRec: A New Dataset for Multimodal Intent RecognitionHanlei Zhang, Hua Xu, Xin Wang, Qianrui Zhou et al.ACM MM 2022 · 66 citations
- Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionQianrui Zhou, Hua Xu, Hao Li, Hanlei Zhang et al.AAAI 2024 · 45 citations
Related papers
- Contextual Augmented Global Contrast for Multimodal Intent RecognitionKaili Sun, Zhiwen Xie, Mang Ye, Huyin ZhangCVPR 2024 · 19 citations
- M2Lens: Visualizing and Explaining Multimodal Models for Sentiment AnalysisXingbo Wang, Jianben He, Zhihua Jin, Muqiao Yang et al.IEEE VIS 2021 · 5 citations
- Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent RecognitionBo Hu, Kai Zhang, Yanghai Zhang, Yuyang YeAAAI 2025 · 6 citations
- InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory ProcessingZhihong Zhu, Xuxin Cheng, Zhaorun Chen, Yuyan Chen et al.ACM MM 2024 · 11 citations
- CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment AnalysisHaoyu Jiang, Xiaoliang Chen, Duoqian Miao, Xiaolin Qin et al.CVPR 2026
