InMu-Net: Advancing Multi-modal Intent Detection via Information Bottleneck and Multi-sensory Processing
Zhihong Zhu, Xuxin Cheng, Zhaorun Chen, Yuyan Chen, Yunyan Zhang, Xian Wu, Yefeng Zheng, Bowen Xing
摘要
Multi-modal intent detection (MID) aims to comprehend users' intentions through diverse modalities, which has received widespread attention in dialogue systems. Despite the promising advancements in complex fusion mechanisms or architecture designs, challenges remain due to: (1) various noise and redundancy in both visual and audio modalities and (2) long-tailed distributions of intent categories. In this paper, to tackle the above two issues, we propose InMu-Net, a simple yet effective framework for MID from the Information bottleneck and Multi-sensory processing perspective. Our contributions lie in three aspects. First, we devise a denoising bottleneck module to filter out the intent-irrelevant information in the fused feature; Second, we introduce a saliency preservation loss to prevent the dropping of intent-relevant information; Ultimately, kurtosis regulation is introduced to maintain representation smoothness during the filtering process, mitigating the adverse impact of the long tail distribution. Comprehensive experiments on two MID benchmark datasets demonstrate the effectiveness of InMu-Net and its vital components. Impressively, a series of analyses reveal our denoising potential and robustness in low-resource, modality corruption, cross-architecture and cross-task scenarios.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper5
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 被引用 17 次
- Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent RecognitionQianrui Zhou, Hua Xu, Yunjin Gu, Yifan Wang 等CVPR 2026 · 被引用 3 次
- Who Should I Trust? Explicit Confidence-Focused Multimodal Intent RecognitionYi Liu, Qimeng Yang, Lanlan LuAAAI 2026
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent RecognitionQianrui Zhou, Hua Xu, Yifan Wang, Xinzhi Dong 等EMNLP 2025
- SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent RecognitionQin Li, Wenbo Zhang, Limei Liu, Han Peng 等CVPR 2026
相关 Paper
- Denoising Bottleneck with Mutual Information Maximization for Video Multimodal FusionShaoxiang Wu, Damai Dai, Ziwei Qin, Tianyu Liu 等ACL 2023 · 被引用 11 次
- Dual-oriented Disentangled Network with Counterfactual Intervention for Multimodal Intent DetectionZhanpeng Chen, Zhihong Zhu, Xianwei Zhuang, Zhiqi Huang 等EMNLP 2024 · 被引用 4 次
- Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck PerspectiveKaifang Long, Lianbo Ma, Jiaqi Liu, liming liu 等CVPR 2026 · 被引用 5 次
- DiffuFuse: Diffusion-Driven Dual-Stream Fusion Framework for Multimodal Sentiment AnalysisXiongjian Lv, Yimin Wen, Hang YuACM MM 2025
- Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal LearningXiongye Xiao, Gengshuo Liu, Gaurav Gupta, Defu Cao 等ICLR 2024 · 被引用 32 次
