Contextual Augmented Global Contrast for Multimodal Intent Recognition
Kaili Sun, Zhiwen Xie, Mang Ye, Huyin Zhang
摘要
Multimodal intent recognition (MIR) aims to perceive the human intent polarity via language, visual, and acoustic modalities. The inherent intent ambiguity makes it challenging to recognize in multimodal scenarios. Existing MIR methods tend to model the individual video independently, ignoring global contextual information across videos. This learning manner inevitably introduces perception biases, exacerbated by the inconsistencies of the multimodal representation, amplifying the intent uncertainty. This challenge motivates us to explore effective global context modeling. Thus, we propose a context-augmented global contrast (CAGC) method to capture rich global context features by mining both intra-and cross-video context interactions for MIR. Concretely, we design a context-augmented transformer module to extract global context dependencies across videos. To further alleviate error accumulation and interference, we develop a cross-video bank that retrieves effective video sources by considering both intentional tendency and video similarity. Furthermore, we introduce a global context-guided contrastive learning scheme, designed to mitigate inconsistencies arising from global context and individual modalities in different feature spaces. This scheme incorporates global cues as the supervision to capture robust the multimodal intent representation. Experiments demonstrate CAGC obtains superior performance than state-of-the-art MIR methods. We also generalize our approach to a closely related task, multimodal sentiment analysis, achieving the comparable performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu 等CVPR 2026 · 被引用 7 次
- Adaptive Multimodal Fusion: Dynamic Attention Allocation for Intent RecognitionBo Hu, Kai Zhang, Yanghai Zhang, Yuyang YeAAAI 2025 · 被引用 6 次
- TiCAL: Typicality-Based Consistency-Aware Learning for Multimodal Emotion RecognitionWen Yin, Siyu Zhan, Cencen Liu, Xin Hu 等AAAI 2026 · 被引用 4 次
- Adaptive Re-calibration Learning for Balanced Multimodal Intention RecognitionQu Yang, Xiyang Li, Fu Lin, Mang YeNeurIPS 2025 · 被引用 2 次
- Group Cognition Learning: Making Everything Better Through Controlled Two-Stage Agents CollaborationChunlei Meng, Pengbin Feng, Rong Fu, Hoi Leong Lee 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper33
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
相关 Paper
- Token-Level Contrastive Learning with Modality-Aware Prompting for Multimodal Intent RecognitionQianrui Zhou, Hua Xu, Hao Li, Hanlei Zhang 等AAAI 2024 · 被引用 45 次
- ConFEDE: Contrastive Feature Decomposition for Multimodal Sentiment AnalysisJiuding Yang, Yakun Yu, Di Niu, Weidong Guo 等ACL 2023 · 被引用 135 次
- Video Entailment via Reaching a Structure-Aware Cross-modal ConsensusXuan Yao, Junyu Gao, Mengyuan Chen, Changsheng XuACM MM 2023 · 被引用 4 次
- Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisHaoyu Zhang, Yu Wang, Guanghao Yin, Kejun Liu 等EMNLP 2023 · 被引用 131 次
- Accommodating Audio Modality in CLIP for Multimodal ProcessingLudan Ruan, Anwen Hu, Yuqing Song, Liang Zhang 等AAAI 2023 · 被引用 18 次
