Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
Zhiyuan Zhou, Yanrong Guo, Shijie Hao
Abstract
Emotional and cognitive factors are essential for understanding mental health disorders. However, existing methods often treat multi-modal data as classification tasks, limiting interpretability especially for emotion and cognition. Although large language models (LLMs) offer opportunities for mental health analysis, they mainly rely on textual semantics and overlook fine-grained emotional and cognitive cues in multi-modal inputs. While some studies incorporate emotional features via transfer learning, their connection to mental health conditions remains implicit. To address these issues, we propose ECMC, a novel task that aims at generating natural language descriptions of emotional and cognitive states from multi-modal data, and producing emotion–cognition profiles that improve both the accuracy and interpretability of mental health assessments. We adopt an encoder–decoder architecture, where modality-specific encoders extract features, which are fused by a dual-stream BridgeNet based on Q-former. Contrastive learning enhances the extraction of emotional and cognitive features. A LLaMA decoder then aligns these features with annotated captions to produce detailed descriptions. Extensive objective and subjective evaluations demonstrate that: 1) ECMC outperforms existing multi-modal LLMs and mental health models in generating emotion–cognition captions; 2) the generated emotion–cognition profiles significantly improve assistive diagnosis and interpretability in mental health analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b68f5de-1d8d-4c83-8581-55df6cf1a618Builds on6
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 2,336 citations
- Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction TuningZebang Cheng, Zhi-Qi Cheng, Jun-Yan He, Kai Wang et al.NeurIPS 2024 · 293 citations
- SECap: Speech Emotion Captioning with Large Language ModelYaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang et al.AAAI 2024 · 70 citations
- Dual-path Collaborative Generation Network for Emotional Video CaptioningCheng Ye, Weidong Chen, Jingyu Li, Lei Zhang et al.ACM MM 2024 · 15 citations
Related papers
- Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion DetectionJun Seo Kim, Hyemi Kim, Woo Joo Oh, Hongjin Cho et al.ACL 2026
- LENS: LLM-Enabled Narrative Synthesis for Mental Health by Aligning Multimodal Sensing with Language ModelsWenxuan Xu, Arvind Pillai, Subigya Nepal, Amanda C. Collins et al.ACL 2026 · 1 citation
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language ModelsZachary Englhardt, Chengqian Ma, Margaret E. Morris, Chun-Cheng Chang et al.UbiComp 2024 · 51 citations
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie et al.EMNLP 2023 · 114 citations
- Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support ConversationLin Zhong, Renjin Zhu, Shujuan Ma, Jinhao Cui et al.ACL 2026
