Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment Detection
Yiwei Wei, Shaozu Yuan, Ruosong Yang, Lei Shen, Zhangmeizhi Li, Longbiao Wang, Meng Chen
摘要
With the popularity of social media, detecting sentiment from multimodal posts (e.g. imagetext pairs) has attracted substantial attention recently. Existing works mainly focus on fusing different features but ignore the challenge of modality heterogeneity. Specifically, different modalities with inherent disparities may bring three problems: 1) introducing redundant visual features during feature fusion; 2) causing feature shift in the representation space; 3) leading to inconsistent annotations for different modal data. All these issues will increase the difficulty in understanding the sentiment of the multimodal content. In this paper, we propose a novel Multi-View Calibration Network (MVCN) to alleviate the above issues systematically. We first propose a text-guided fusion module with novel Sparse-Attention to reduce the negative impacts of redundant visual elements. We then devise a sentiment-based congruity constraint task to calibrate the feature shift in the representation space. Finally, we introduce an adaptive loss calibration strategy to tackle inconsistent annotated labels. Extensive experiments demonstrate the competitiveness of MVCN against previous approaches and achieve state-of-the-art results on two public benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- G^2SAM: Graph-Based Global Semantic Awareness Method for Multimodal Sarcasm DetectionYiwei Wei, Shaozu Yuan, Hengyang Zhou, Longbiao Wang 等AAAI 2024 · 被引用 30 次
- Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and FusionDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 被引用 13 次
- D2R: Dual-Branch Dynamic Routing Network for Multimodal Sentiment DetectionYifan Chen, Kuntao Li, Weixing Mai, Qiaofeng Wu 等EMNLP 2024 · 被引用 10 次
- Bridging Visual Affective Gap: Borrowing Textual Knowledge by Learning from Noisy Image-Text PairsDaiqing Wu, Dongbao Yang, Yu Zhou, Can MaACM MM 2024 · 被引用 6 次
- Plug-and-play Feature Causality Decomposition for Multimodal Representation LearningYe Liu, Zihan Ji, Hongmin CaiNeurIPS 2025 · 被引用 4 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty 等NeurIPS 2021 · 被引用 2,985 次
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung 等NeurIPS 2022 · 被引用 834 次
相关 Paper
- Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution ConstraintsQinfu Xu, Yiwei Wei, Chunlei Wu, Leiquan Wang 等AAAI 2025 · 被引用 3 次
- DealMVC: Dual Contrastive Calibration for Multi-view ClusteringXihong Yang, Jiaqi Jin, Siwei Wang, Ke Liang 等ACM MM 2023 · 被引用 138 次
- CTFN: Hierarchical Learning for Multimodal Sentiment Analysis Using Coupled-Translation Fusion NetworkJiajia Tang, Kang Li, Xuanyu Jin, Andrzej Cichocki 等ACL 2021
- Multi-Modal Sarcasm Detection with Interactive In-Modal and Cross-Modal GraphsBin Liang, Chenwei Lou, Xiang Li, Lin Gui 等ACM MM 2021 · 被引用 128 次
- Multispectral Object Detection via Cross-Modal Conflict-Aware LearningXiao He, Chang Tang, Xin Zou, Wei ZhangACM MM 2023 · 被引用 84 次
