Multimodal fusion via cortical network inspired losses
Shiv Shankar
Abstract
Information integration from different modalities is an active area of research. Human beings and, in general, biological neural systems are quite adept at using a multitude of signals from different sensory perceptive fields to interact with the environment and each other. Recent work in deep fusion models via neural networks has led to substantial improvements over unimodal approaches in areas like speech recognition, emotion recognition and analysis, captioning and image description. However, such research has mostly focused on architectural changes allowing for fusion of different modalities while keeping the model complexity manageable. Inspired by neuroscientific ideas about multisensory integration and processing, we investigate the effect of introducing neural dependencies in the loss functions. Experiments on multimodal sentiment analysis tasks with different models show that our approach provides a consistent performance boost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f78452b4-2c51-4f19-9c44-7478028af403Cited by top-tier papers5
- Joyful: Joint Modality Fusion and Graph Contrastive Learning for Multimoda Emotion RecognitionDongyuan Li, Yusong Wang, Kotaro Funakoshi, Manabu OkumuraEMNLP 2023 · 46 citations
- Curriculum Learning Meets Weakly Supervised Multimodal Correlation LearningSijie Mai, Ya Sun, Haifeng HuEMNLP 2022 · 9 citations
- On Online Experimentation without Device IdentifiersShiv Shankar, Ritwik Sinha, Madalina FiterauICML 2024 · 2 citations
- CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal FusionSijie Mai, Shiqin HanCVPR 2026 · 1 citation
- Learning Invariant Modality Representation for Robust Multimodal Learning from a Causal Inference PerspectiveSijie Mai, Shiqin HanACL 2026
Builds on9
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Learning Modality-Specific Representations with Self-Supervised Multi-Task Learning for Multimodal Sentiment AnalysisWenmeng Yu, Hua Xu, Ziqi Yuan, Jiele WuAAAI 2021 · 737 citations
- Integrating Multimodal Information in Large Pretrained TransformersWasifur Rahman, Md. Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh et al.ACL 2020 · 584 citations
- Episodic Training for Domain GeneralizationDa Li, Jianshu Zhang, Yongxin Yang, Cong Liu et al.ICCV 2019 · 488 citations
- Learning Relationships between Text, Audio, and Video via Deep Canonical Correlation for Multimodal Language AnalysisZhongkai Sun, Prathusha Kameswara Sarma, William A. Sethares, Yingyu LiangAAAI 2020 · 419 citations
Related papers
- Improving Multimodal fusion via Mutual Dependency MaximisationPierre Colombo, Emile Chapuis, Matthieu Labeau, Chloé ClavelEMNLP 2021 · 29 citations
- Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment AnalysisWei Han, Hui Chen, Soujanya PoriaEMNLP 2021 · 9 citations
- Factorize, Reconstruct, Enhance: A Unified Framework for Multimodal Sentiment AnalysisZhilu Yang, Mingcheng LiCVPR 2026
- PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment AnalysisHeng Xie, Kang Zhu, Zhengqi Wen, Jianhua Tao et al.AAAI 2026 · 1 citation
- ND-MRM: Neuronal Diversity Inspired Multisensory Recognition ModelQixin Wang, Chaoqiong Fan, Tianyuan Jia, Yuyang Han et al.AAAI 2024 · 5 citations
