Dep-MAP: A Multi-level Alignment Framework with Semantic Prototypes for Video-based Automatic Depression Assessment
Hao Wang, Jiayu Ye, Qingxiang Wang
Abstract
Spatiotemporal analysis of facial behavior is a crucial method for evaluating the mental state of depression patients. However, in practice, depressed patients often display facial behaviors similar to healthy individuals due to masking tendencies. Additionally, facial expressions among depressed patients are also different, increasing the difficulty of assessment. To address this, we propose a video-based automatic depression assessment model Dep-MAP for complex facial behaviors of depression patients. Dep-MAP adopts a dual-branch architecture to extract visual features of facial behavior and capture corresponding emotional semantic features. Specifically, the extracted deep semantic features are clustered, resulting in semantically distinct prototype sets, where each severity group learns a set of discriminative facial behavior prototype representations, to suppress inter-class semantic confusion. Subsequently, we propose a semantic prototype-supervised contrastive learning method, which aligns latent semantics between shallow and deep features, realizing emotional semantic guidance and self-knowledge distillation for the visual feature branch, effectively suppressing intra-class difference. Then, we integrate key depression cues across multiple spatiotemporal scales via a multi-scale weighted fusion strategy, achieving automatic depression assessment. Experimental results demonstrate that Dep-MAP effectively identifies potential key frames in temporal sequences, and aggregates key frame representations with semantic consistency, achieving significantly superior state-of-the-art results on the AVEC2013 and AVEC2014 public datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01a4d2b7-d650-47cc-b843-91466bb0dff9Builds on4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression AssessmentZijian Wu, Leijing Zhou, Shuanglin Li, Changzeng Fu et al.AAAI 2025 · 6 citations
- Feature Decomposition and Reconstruction Learning for Effective Facial Expression RecognitionDelian Ruan, Yan Yan, Shenqi Lai, Zhenhua Chai et al.CVPR 2021
- Exploiting Semantic Embedding and Visual Feature for Facial Action Unit DetectionHuiyuan Yang, Lijun Yin, Yi Zhou, Jiuxiang GuCVPR 2021
Related papers
- Explainable Depression Assessment from Face Videos by Weakly Supervised LearningRongfan Liao, Xiangyu Kong, Shiqing Tang, Lang He et al.AAAI 2026
- FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksRuiqi Wang, Jinyang Huang, Jie Zhang, Xin Liu et al.ACM MM 2024 · 22 citations
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 11 citations
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj et al.ICCV 2021 · 76 citations
- Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior UnderstandingXiang Zhang, Taoyue Wang, Xiaotian Li, Huiyuan Yang et al.ICCV 2023 · 26 citations
