Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution Matching
Jingjun Liang, Ruichen Li, Qin Jin
Abstract
Automatic emotion recognition is an active research topic with wide range of applications. Due to the high manual annotation cost and inevitable label ambiguity, the development of emotion recognition dataset is limited in both scale and quality. Therefore, one of the key challenges is how to build effective models with limited data resource. Previous works have explored different approaches to tackle this challenge including data enhancement, transfer learning, and semi-supervised learning etc. However, the weakness of these existing approaches includes such as training instability, large performance loss during transfer, or marginal improvement. In this work, we propose a novel semi-supervised multi-modal emotion recognition model based on cross-modality distribution matching, which leverages abundant unlabeled data to enhance the model training under the assumption that the inner emotional status is consistent at the utterance level across modalities. We conduct extensive experiments to evaluate the proposed model on two benchmark datasets, IEMOCAP and MELD. The experiment results prove that the proposed semi-supervised learning model can effectively utilize unlabeled data and combine multi-modalities to boost the emotion recognition performance, which outperforms other state-of-the-art approaches under the same condition. The proposed model also achieves competitive capacity compared with existing approaches which take advantage of additional auxiliary information such as speaker and interaction context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dac9c733-50f1-47d2-bce6-beb0abedb8c4Cited by top-tier papers8
- Towards Robust Multimodal Sentiment Analysis with Incomplete DataHaoyu Zhang, Wenbin Wang, Tianshu YuNeurIPS 2024 · 90 citations
- DocTr: Document Image Transformer for Geometric Unwarping and Illumination CorrectionHao Feng, Yuechen Wang, Wengang Zhou, Jiajun Deng et al.ACM MM 2021 · 66 citations
- A Facial Expression-Aware Multimodal Multi-task Learning Framework for Emotion Recognition in Multi-party ConversationsWenjie Zheng, Jianfei Yu, Rui Xia, Shijin WangACL 2023 · 38 citations
- Semi-IIN: Semi-Supervised Intra-Inter Modal Interaction Learning Network for Multimodal Sentiment AnalysisJinhao Lin, Yifei Wang, Yanwu Xu, Qi LiuAAAI 2025 · 3 citations
- Beyond Static Alignment: Adaptive Arbitration for Semantic Incongruence in Semi-Supervised Multimodal Sentiment AnalysisHuicong Li, Xiangbo Ji, Wei WuACL 2026
Related papers
- Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning ApproachSuping Zhou, Jia Jia, Zhiyong Wu, Zhihan Yang et al.AAAI 2021 · 20 citations
- EMOE: Modality-Specific Enhanced Dynamic Emotion ExpertsYiyang Fang, Wenke Huang, Guancheng Wan, Kehua Su et al.CVPR 2025
- Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift ModelingXupeng Zha, Huan Zhao, Guanghui Ye, Zixing ZhangAAAI 2025 · 3 citations
- MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in ConversationJingwen Hu, Yuchen Liu, Jinming Zhao, Qin JinACL 2021
- HeLo: Heterogeneous Multi-Modal Fusion with Label Correlation for Emotion Distribution LearningChuhang Zheng, Chunwei Tian, Jie Wen, Daoqiang Zhang et al.ACM MM 2025 · 13 citations
