Multi-label Emotion Analysis in Conversation via Multimodal Knowledge Distillation
Sidharth Anand, Naresh Kumar Devulapally, Sreyasee Das Bhattacharjee, Junsong Yuan
Abstract
Evaluating speaker emotion in conversations is crucial for various applications requiring human-computer interaction. However, co-occurrences of multiple emotional states (e.g. 'anger' and 'frustration' may occur together or one may influence the occurrence of the other) and their dynamic evolution may vary dramatically due to the speaker's internal (e.g., influence of their personalized socio-cultural-educational and demographic backgrounds) and external contexts. Thus far, the previous focus has been on evaluating only the dominant emotion observed in a speaker at a given time, which is susceptible to producing misleading classification decisions for difficult multi-labels during testing. In this work, we present Self-supervised Multi- Label Peer Collaborative Distillation (SeMuL-PCD) Learning via an efficient Multimodal Transformer Network, in which complementary feedback from multiple mode-specific peer networks (e.g.transcript, audio, visual) are distilled into a single mode-ensembled fusion network for estimating multiple emotions simultaneously. The proposed Multimodal Distillation Loss calibrates the fusion network by minimizing the Kullback-Leibler divergence with the peer networks. Additionally, each peer network is conditioned using a self-supervised contrastive objective to improve the generalization across diverse socio-demographic speaker backgrounds. By enabling peer collaborative learning that allows each network to independently learn their mode-specific discriminative patterns,SeMUL-PCD is effective across different conversation environments. In particular, the model not only outperforms the current state-of-the-art models on several large-scale public datasets (e.g., MOSEI, EmoReact and ElderReact), but with around 17% improved weighted F1-score in the cross-dataset experimental settings. The model also demonstrates an impressive generalization ability across age and demography-diverse populations.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 1c472650-c7b3-4ead-b9e9-2dadda969c5cCited by top-tier papers2
- From Individuals to Crowds: Dual-Level Public Response Prediction in Social MediaJinghui Zhang, Kaiyang Wan, Longwei Xu, Ao Li et al.ACM MM 2025
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive LearningJiangfeng Sun, Sihao He, Zhonghong Ou, Meina SongAAAI 2026
Related papers
- A Unimodal Valence-Arousal Driven Contrastive Learning Framework for Multimodal Multi-Label Emotion RecognitionWenjie Zheng, Jianfei Yu, Rui XiaACM MM 2024 · 8 citations
- Multimodal Prompt Transformer with Hybrid Contrastive Learning for Emotion Recognition in ConversationShihao Zou, Xianying Huang, Xudong ShenACM MM 2023 · 24 citations
- CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing ModalitiesYan Zhuang, Minhao Liu, Wei Bai, Yanru Zhang et al.ICCV 2025 · 2 citations
- Semi-supervised Multi-modal Emotion Recognition with Cross-Modal Distribution MatchingJingjun Liang, Ruichen Li, Qin JinACM MM 2020 · 67 citations
- UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionGuimin Hu, Ting-En Lin, Yi Zhao, Guangming Lu et al.EMNLP 2022 · 206 citations
