Dep-MAP: A Multi-level Alignment Framework with Semantic Prototypes for Video-based Automatic Depression Assessment
Hao Wang, Jiayu Ye, Qingxiang Wang
摘要
Spatiotemporal analysis of facial behavior is a crucial method for evaluating the mental state of depression patients. However, in practice, depressed patients often display facial behaviors similar to healthy individuals due to masking tendencies. Additionally, facial expressions among depressed patients are also different, increasing the difficulty of assessment. To address this, we propose a video-based automatic depression assessment model Dep-MAP for complex facial behaviors of depression patients. Dep-MAP adopts a dual-branch architecture to extract visual features of facial behavior and capture corresponding emotional semantic features. Specifically, the extracted deep semantic features are clustered, resulting in semantically distinct prototype sets, where each severity group learns a set of discriminative facial behavior prototype representations, to suppress inter-class semantic confusion. Subsequently, we propose a semantic prototype-supervised contrastive learning method, which aligns latent semantics between shallow and deep features, realizing emotional semantic guidance and self-knowledge distillation for the visual feature branch, effectively suppressing intra-class difference. Then, we integrate key depression cues across multiple spatiotemporal scales via a multi-scale weighted fusion strategy, achieving automatic depression assessment. Experimental results demonstrate that Dep-MAP effectively identifies potential key frames in temporal sequences, and aggregates key frame representations with semantic consistency, achieving significantly superior state-of-the-art results on the AVEC2013 and AVEC2014 public datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- DepMGNN: Matrixial Graph Neural Network for Video-based Automatic Depression AssessmentZijian Wu, Leijing Zhou, Shuanglin Li, Changzeng Fu 等AAAI 2025 · 被引用 6 次
- Feature Decomposition and Reconstruction Learning for Effective Facial Expression RecognitionDelian Ruan, Yan Yan, Shenqi Lai, Zhenhua Chai 等CVPR 2021
- Exploiting Semantic Embedding and Visual Feature for Facial Action Unit DetectionHuiyuan Yang, Lijun Yin, Yi Zhou, Jiuxiang GuCVPR 2021
相关 Paper
- Explainable Depression Assessment from Face Videos by Weakly Supervised LearningRongfan Liao, Xiangyu Kong, Shiqing Tang, Lang He 等AAAI 2026
- FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksRuiqi Wang, Jinyang Huang, Jie Zhang, Xin Liu 等ACM MM 2024 · 被引用 22 次
- MART: Masked Affective RepresenTation Learning via Masked Temporal Distribution DistillationZhicheng Zhang, Pancheng Zhao, Eunil Park, Jufeng YangCVPR 2024 · 被引用 11 次
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj 等ICCV 2021 · 被引用 76 次
- Weakly-Supervised Text-driven Contrastive Learning for Facial Behavior UnderstandingXiang Zhang, Taoyue Wang, Xiaotian Li, Huiyuan Yang 等ICCV 2023 · 被引用 26 次
