Decoupled Multimodal Distilling for Emotion Recognition
Yong Li, Yuanzhi Wang, Zhen Cui
Abstract
Human multimodal emotion recognition (MER) aims to perceive human emotions via language, visual and acoustic modalities. Despite the impressive performance of previous MER approaches, the inherent multimodal heterogeneities still haunt and the contribution of different modalities varies significantly. In this work, we mitigate this issue by proposing a decoupled multimodal distillation (DMD) approach that facilitates flexible and adaptive crossmodal knowledge distillation, aiming to enhance the discriminative features of each modality. Specially, the representation of each modality is decoupled into two parts, i.e., modalityirrelevant/-exclusive spaces, in a self-regression manner. DMD utilizes a graph distillation unit (GD-Unit) for each decoupled part so that each GD can be performed in a more specialized and effective manner. A GD-Unit consists of a dynamic graph where each vertice represents a modality and each edge indicates a dynamic knowledge distillation. Such GD paradigm provides a flexible knowledge transfer manner where the distillation weights can be automatically learned, thus enabling diverse crossmodal knowledge transfer patterns. Experimental results show DMD consistently obtains superior performance than state-of-the-art MER methods. Visualization results show the graph edges in DMD exhibit meaningful distributional patterns w.r.t. the modality-irrelevant/-exclusive feature spaces. Codes are released at https://github.com/mdswyz/DMD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02f384a7-5467-41ef-a6e7-74a028822577Cited by top-tier papers45
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang et al.NeurIPS 2024 · 293 citations
- Incomplete Multimodality-Diffused Emotion RecognitionYuanzhi Wang, Yong Li, Zhen CuiNeurIPS 2023 · 155 citations
- Distribution-Consistent Modal Recovering for Incomplete Multimodal LearningYuanzhi Wang, Zhen Cui, Yong LiICCV 2023 · 101 citations
- DLF: Disentangled-Language-Focused Multimodal Sentiment AnalysisPan Wang, Qiang Zhou, Yawen Wu, Tianlong Chen et al.AAAI 2025 · 84 citations
- A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing ModalitiesMingcheng Li, Dingkang Yang, Yuxuan Lei, Shunli Wang et al.AAAI 2024 · 71 citations
Builds on9
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu et al.ICML 2022 · 1,123 citations
- MISA: Modality-Invariant and -Specific Representations for Multimodal Sentiment AnalysisDevamanyu Hazarika, Roger Zimmermann, Soujanya PoriaACM MM 2020 · 1,037 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
Related papers
- EMOE: Modality-Specific Enhanced Dynamic Emotion ExpertsYiyang Fang, Wenke Huang, Guancheng Wan, Kehua Su et al.CVPR 2025
- EmotionKD: A Cross-Modal Knowledge Distillation Framework for Emotion Recognition Based on Physiological SignalsYucheng Liu, Ziyu Jia, Haichao WangACM MM 2023 · 55 citations
- DRKF: Decoupled Representations with Knowledge Fusion for Multimodal Emotion RecognitionPeiyuan Jiang, Yao Liu, Qiao Liu, Zongshun Zhang et al.ACM MM 2025 · 4 citations
- Disentangled-Multimodal Privileged Knowledge Distillation for Depression Recognition with Incomplete Multimodal DataYuchen Pan, Junjun Jiang, Kui Jiang, Xianming LiuACM MM 2024 · 18 citations
- Correlation-Driven Multi-Modality Graph Decomposition for Cross-Subject Emotion RecognitionWuliang Huang, Yiqiang Chen, Xinlong Jiang, Chenlong Gao et al.ACM MM 2024 · 2 citations
