Improving Multimodal fusion via Mutual Dependency Maximisation
Pierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé Clavel
摘要
Multimodal sentiment analysis is a trending area of research, and the multimodal fusion is one of its most active topic. Acknowledging humans communicate through a variety of channels (i.e visual, acoustic, linguistic), multimodal systems aim at integrating different unimodal representations into a synthetic one. So far, a consequent effort has been made on developing complex architectures allowing the fusion of these modalities. However, such systems are mainly trained by minimising simple losses such as L 1 or cross-entropy. In this work, we investigate unexplored penalties and propose a set of new objectives that measure the dependency between modalities. We demonstrate that our new penalties lead to a consistent improvement (up to 4.3 on accuracy) across a large variety of state-of-the-art models on two well-known sentiment analysis datasets: CMU-MOSI and CMU-MOSEI. Our method not only achieves a new SOTA on both datasets but also produces representations that are more robust to modality drops. Finally, a by-product of our methods includes a statistical network which can be used to interpret the high dimensional representations learnt by the model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Multimodal Variational Auto-encoder based Audio-Visual SegmentationYuxin Mao, Jing Zhang, Mochu Xiang, Yiran Zhong 等ICCV 2023 · 被引用 57 次
- Beyond Mahalanobis Distance for Textual OOD DetectionPierre Colombo, Eduardo Dadalto Câmara Gomes, Guillaume Staerman, Nathan Noiry 等NeurIPS 2022 · 被引用 24 次
- Multimodal fusion via cortical network inspired lossesShiv ShankarACL 2022 · 被引用 21 次
- What are the best Systems? New Perspectives on NLP BenchmarkingPierre Colombo, Nathan Noiry, Ekhine Irurozki, Stéphan ClémençonNeurIPS 2022 · 被引用 20 次
- MM-Align: Learning Optimal Transport-based Alignment Dynamics for Fast and Accurate Inference on Missing Modality SequencesWei Han, Hui Chen, Min-Yen Kan, Soujanya PoriaEMNLP 2022 · 被引用 13 次
它引用的顶会 Paper7
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 被引用 419 次
- Understanding the Limitations of Variational Mutual Information EstimatorsJiaming Song, Stefano ErmonICLR 2020 · 被引用 243 次
- A Mutual Information Maximization Perspective of Language Representation LearningLingpeng Kong, Cyprien de Masson d'Autume, Lei Yu, Wang Ling 等ICLR 2020 · 被引用 179 次
相关 Paper
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive LearningJiangfeng Sun, Sihao He, Zhonghong Ou, Meina SongAAAI 2026
- Tri-Subspaces Disentanglement for Multimodal Sentiment AnalysisChunlei Meng, Jiabin Luo, Zhenglin Yan, Zhenyu Yu 等CVPR 2026 · 被引用 7 次
- DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisPan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen 等AAAI 2025 · 被引用 84 次
- Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment AnalysisZhilu Yang, Mingcheng LiCVPR 2026
- CICA: Coupling Confidence-Aware Pretraining with Confidence-Informed Attention for Robust Multimodal Sentiment AnalysisHaoyu Jiang, Xiaoliang Chen, Duoqian Miao, Xiaolin Qin 等CVPR 2026
