MDDR: Multi-modal Dual-Attention aggregation for Depression Recognition
Wei Zhang, En Zhu, Juan Chen, Yunpeng Li
摘要
Automated diagnosis of depression is crucial for early detection and timely intervention. Previous research has largely concentrated on visual information, often neglecting the value of leveraging a variety of data types. Although some studies have attempted to employ multiple modalities, they typically fall short in investigating the complex dynamics between features from various modalities over time. To address this challenge, we present an innovative Multi-modal Dual-Attention aggregation architecture for Depression Recognition (MDDR). This framework leverages multi-modal pre-trained features and introduces two attention aggregation mechanisms: the Feature Alignment and Aggregation (FAA) module and the Sequence Encoding and Aggregation (SEA) module. The FAA module is designed to dynamically evaluate the relevance of multi-modal features for each instance, facilitating a dynamic integration of these features over time. Following this, the SEA module determines the importance of the amalgamated features for each frame, ensuring that aggregation is conducted based on their significance, to extract the most relevant features for accurately diagnosing depression. Moreover, we propose a unique loss calculation method specifically designed for depression assessment, named DR Loss. Our approach, evaluated on the AVEC2013 and AVEC2014 depression audiovisual datasets, achieves unparalleled performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Disentangled-Multimodal Privileged Knowledge Distillation for Depression Recognition with Incomplete Multimodal DataYuchen Pan, Junjun Jiang, Kui Jiang, Xianming LiuACM MM 2024 · 被引用 18 次
- Explainable Depression Assessment from Face Videos by Weakly Supervised LearningRongfan Liao, Xiangyu Kong, Shiqing Tang, Lang He 等AAAI 2026
- FacialPulse: An Efficient RNN-based Depression Detection via Temporal Facial LandmarksRuiqi Wang, Jinyang Huang, Jie Zhang, Xin Liu 等ACM MM 2024 · 被引用 22 次
- Dep-MAP: A Multi-level Alignment Framework with Semantic Prototypes for Video-based Automatic Depression AssessmentHao Wang, Jiayu Ye, Qingxiang WangAAAI 2026
- A Multimodal EEG-Eye Movement Model for Automatic Depression DetectionHao-Long Yin, Jian-Ming Zhang, Ren-Jie Dai, Wei-Long Zheng 等AAAI 2026
