Disentangled-Multimodal Privileged Knowledge Distillation for Depression Recognition with Incomplete Multimodal Data
Yuchen Pan, Junjun Jiang, Kui Jiang, Xianming Liu
Abstract
Depression recognition (DR) using facial images, audio signals, or language text recordings has achieved remarkable performance. Recently, multimodal DR has shown improved performance over single-modal methods by leveraging information from a combination of these modalities. However, collecting high-quality data containing all modalities poses a challenge. In particular, these methods often encounter performance degradation when certain modalities are either missing or degraded. To tackle this issue, we present a generalizable multimodal framework for DR by aggregating feature disentanglement and privileged knowledge distillation. In detail, our approach aims to disentangle homogeneous and heterogeneous features within multimodal signals while suppressing noise, thereby adaptively aggregating the most informative components for high-quality DR. Subsequently, we leverage knowledge distillation to transfer privileged knowledge from complete modalities to the observed input with limited information, thereby significantly improving the tolerance and compatibility. These strategies form our novel Feature Disentanglement and Privileged knowledge Distillation Network for DR, dubbed Dis2DR. Experimental evaluations on AVEC 2013, AVEC 2014, AVEC 2017, and AVEC 2019 datasets demonstrate the effectiveness of our Dis2DR method. Remarkably, Dis2DR achieves superior performance even when only a single modality is available, surpassing existing state-of-the-art multimodal DR approaches AVA-DepressNet by up to 9.8% on the AVEC 2013 dataset.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- Asymmetric Reinforcing Against Multi-Modal Representation BiasXiyuan Gao, Bing Cao, Pengfei Zhu, Nannan Wang et al.AAAI 2025 · 6 citations
- Memory-Efficient Transfer Learning with Fading Side Networks via Masked Dual Path DistillationYutong Zhang, Jiaxin Chen, Honglin Chen, Kaiqi Zheng et al.CVPR 2026 · 1 citation
Related papers
- MDDR: Multi-modal Dual-Attention aggregation for Depression RecognitionWei Zhang, En Zhu, Juan Chen, Yunpeng LiACM MM 2024 · 8 citations
- Decoupled Multimodal Distilling for Emotion RecognitionYong Li, Yuanzhi Wang, Zhen CuiCVPR 2023
- EmotionKD: A Cross-Modal Knowledge Distillation Framework for Emotion Recognition Based on Physiological SignalsYucheng Liu, Ziyu Jia, Haichao WangACM MM 2023 · 55 citations
- OpticalDR: A Deep Optical Imaging Model for Privacy-Protective Depression RecognitionYuchen Pan, Junjun Jiang, Kui Jiang, Zhihao Wu et al.CVPR 2024 · 10 citations
- Disentangled Representation Learning for Multimodal Emotion RecognitionDingkang Yang, Shuai Huang, Haopeng Kuang, Yangtao Du et al.ACM MM 2022 · 260 citations
