Layer-wise Fusion with Modality Independence Modeling for Multi-modal Emotion Recognition
Jun Sun, Shoukang Han, Yu-Ping Ruan, Xiaoning Zhang, Shu-Kai Zheng, Yulong Liu, Yuxin Huang, Taihao Li
Abstract
Multi-modal emotion recognition has gained increasing attention in recent years due to its widespread applications and the advances in multi-modal learning approaches. However, previous studies primarily focus on developing models that exploit the unification of multiple modalities. In this paper, we propose that maintaining modality independence is beneficial for the model performance. According to this principle, we construct a dataset, and devise a multimodal transformer model. The new dataset, CHinese Emotion Recognition dataset with Modality-wise Annotations, abbreviated as CHERMA, provides uni-modal labels for each individual modality, and multi-modal labels for all modalities jointly observed. The model consists of uni-modal transformer modules that learn representations for each modality, and a multi-modal transformer module that fuses all modalities. All the modules are supervised by their corresponding labels separately, and the forward information flow is uni-directionally from the uni-modal modules to the multimodal module. The supervision strategy and the model architecture guarantee each individual modality learns its representation independently, and meanwhile the multimodal module aggregates all information. Extensive empirical results demonstrate that our proposed scheme outperforms state-of-theart alternatives, corroborating the importance of modality independence in multi-modal emotion recognition. The dataset and codes are availabel at https://github.com/ sunjunaimer/LFMIM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc05bf76-7fd7-4fca-b6ee-45898347e341Cited by top-tier papers7
- MSE-Adapter: A Lightweight Plugin Endowing LLMs with the Capability to Perform Multimodal Sentiment Analysis and Emotion RecognitionYang Yang, Xunde Dong, Yupeng QiangAAAI 2025 · 19 citations
- Proxy-Driven Robust Multimodal Sentiment Analysis with Incomplete DataAoqiang Zhu, Min Hu, Xiaohua Wang, Jiaoyun Yang et al.ACL 2025 · 8 citations
- Adversarial Alignment with Anchor Dragging Drift (A³D²): Multimodal Domain Adaptation with Partially Shifted ModalitiesJun Sun, Xinxin Zhang, Simin Hong, Jian Zhu et al.ACL 2025 · 5 citations
- MST-Distill: Mixture of Specialized Teachers for Cross-Modal Knowledge DistillationHui Li, Pengfei Yang, Juanyang Chen, Le Dong et al.ACM MM 2025 · 4 citations
- CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing ModalitiesYan Zhuang, Minhao Liu, Wei Bai, Yanru Zhang et al.ICCV 2025 · 2 citations
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Attention Bottlenecks for Multimodal FusionArsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen et al.NeurIPS 2021 · 884 citations
- CH-SIMS: A Chinese Multimodal Sentiment Analysis Dataset with Fine-grained Annotation of ModalityWenmeng Yu, Hua Xu, Fanyang Meng, Yilin Zhu et al.ACL 2020 · 376 citations
- MISC: A Mixed Strategy-Aware Model integrating COMET for Emotional Support ConversationQuan Tu, Yanran Li, Jianwei Cui, Bin Wang et al.ACL 2022 · 141 citations
- M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue DatabaseJinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu et al.ACL 2022 · 88 citations
Related papers
- Transformer-based Label Set Generation for Multi-modal Multi-label Emotion DetectionXincheng Ju, Dong Zhang, Junhui Li, Guodong ZhouACM MM 2020 · 64 citations
- Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message PassingDong Zhang, Xincheng Ju, Wei Zhang, Junhui Li et al.AAAI 2021 · 56 citations
- Tailor Versatile Multi-Modal Learning for Multi-Label Emotion RecognitionYi Zhang, Mingyuan Chen, Jundong Shen, Chongjun WangAAAI 2022 · 92 citations
- Progressive Modality Reinforcement for Human Multimodal Emotion Recognition From Unaligned Multimodal SequencesFengmao Lv, Xiang Chen, Yanyong Huang, Lixin Duan et al.CVPR 2021
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui et al.CVPR 2023
