Multimodal fusion via cortical network inspired losses
Shiv Shankar
摘要
Information integration from different modalities is an active area of research. Human beings and, in general, biological neural systems are quite adept at using a multitude of signals from different sensory perceptive fields to interact with the environment and each other. Recent work in deep fusion models via neural networks has led to substantial improvements over unimodal approaches in areas like speech recognition, emotion recognition and analysis, captioning and image description. However, such research has mostly focused on architectural changes allowing for fusion of different modalities while keeping the model complexity manageable. Inspired by neuroscientific ideas about multisensory integration and processing, we investigate the effect of introducing neural dependencies in the loss functions. Experiments on multimodal sentiment analysis tasks with different models show that our approach provides a consistent performance boost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 被引用 46 次
- Curriculum Learning Meets Weakly Supervised Multimodal Correlation LearningSijie Mai, Ya Sun, Haifeng HuEMNLP 2022 · 被引用 9 次
- On Online Experimentation without Device IdentifiersShiv Shankar, Ritwik Sinha, Madalina FiterauICML 2024 · 被引用 2 次
- CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal FusionSijie Mai, Shiqin HanCVPR 2026 · 被引用 1 次
- Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference PerspectiveSijie Mai, Shiqin HanACL 2026
它引用的顶会 Paper9
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 被引用 1,037 次
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 被引用 737 次
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh 等ACL 2020 · 被引用 584 次
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu 等ICCV 2019 · 被引用 488 次
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 被引用 419 次
相关 Paper
- Improving Multimodal fusion via Mutual Dependency MaximisationPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 被引用 29 次
- Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment AnalysisWei Han, Hui Chen, Soujanya PoriaEMNLP 2021 · 被引用 9 次
- Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment AnalysisZhilu Yang, Mingcheng LiCVPR 2026
- PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment AnalysisHeng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao 等AAAI 2026 · 被引用 1 次
- ND-MRM: Neuronal Diversity Inspired Multisensory Recognition ModelQixin Wang, Chaoqiong Fan, Tianyuan Jia, Yuyang Han 等AAAI 2024 · 被引用 5 次
