Explainable Depression Assessment from Face Videos by Weakly Supervised Learning
Rongfan Liao, Xiangyu Kong, Shiqing Tang, Lang He, Changzeng Fu, Weicheng Xie, Xiaofeng Liu, Lu Liu, Siyang Song
Abstract
Existing video-based automatic depression assessment (ADA) approaches frequently achieve video-level depression assessment by aggregating features or predictions of individual frames or equal-length segments within the given video. While their performances have been largely enhanced by recent advanced deep learning models, they typically fail to explicitly consider the varied importance of depression-related behavioural cues across different video segments, i.e., segments within one video may contain behaviours reflecting varying levels of depression. Underestimating segment-level variations can obscure the detection of facial behaviour cues associated with depression, thereby undermining the accuracy and interpretability of video-based depression detection systems. In this paper, we propose a novel video-based ADA approach that specifically identifies and differentiates video segments that exhibit depression-related facial behaviours across varying temporal durations, providing clear insights into how each segment contributes to the video-level depression prediction. To achieve this, a novel weakly supervised strategy is proposed to compare segment-level behaviours with video-level depression label, enabling the model to assign depression-relevant scores to multiple temporal scale video segments and attend selectively to those most indicative of depressive states. Extensive experiments on the AVEC 2013 and AVEC 2014 face video depression datasets demonstrate the effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1b13224-5b56-4396-b24c-381a2bb16537Builds on3
- FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksRuiqi Wang, Jinyang Huang, Jie Zhang, Xin Liu et al.ACM MM 2024 · 22 citations
- DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression AssessmentZijian Wu, Leijing Zhou, Shuanglin Li, Changzeng Fu et al.AAAI 2025 · 6 citations
- Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action LocalizationHuan Ren, Wenfei Yang, Tianzhu Zhang, Yongdong ZhangCVPR 2023
Related papers
- Dep-MAP: A Multi-level Alignment Framework with Semantic Prototypes for Video-based Automatic Depression AssessmentHao Wang, Jiayu Ye, Qingxiang WangAAAI 2026
- MDDR: Multi-modal Dual-Attention aggregation for Depression RecognitionWei Zhang, En Zhu, Juan Chen, Yunpeng LiACM MM 2024 · 8 citations
- Disentangled-Multimodal Privileged Knowledge Distillation for Depression Recognition with Incomplete Multimodal DataYuchen Pan, Junjun Jiang, Kui Jiang, Xianming LiuACM MM 2024 · 18 citations
- Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention ReasoningYu Wang, Shengjie ZhaoCVPR 2026 · 6 citations
- Context-Aware Feature and Label Fusion for Facial Action Unit Intensity Estimation With Partially Labeled DataYong Zhang, Haiyong Jiang, Baoyuan Wu, Yanbo Fan et al.ICCV 2019 · 32 citations
