When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
Xiangyu Zhang, Hexin Liu, Kaishuai Xu, Qiquan Zhang, Daijiao Liu, Beena Ahmed, Julien Epps
摘要
Depression is a critical concern in global mental health, prompting extensive research into AIbased detection methods. Among various AI technologies, Large Language Models (LLMs) stand out for their versatility in mental healthcare applications. However, their primary limitation arises from their exclusive dependence on textual input, which constrains their overall capabilities. Furthermore, the utilization of LLMs in identifying and analyzing depressive states is still relatively untapped. In this paper, we present an innovative approach to integrating acoustic speech information into the LLMs framework for multimodal depression detection. We investigate an efficient method for depression detection by integrating speech signals into LLMs utilizing Acoustic Landmarks. By incorporating acoustic landmarks, which are specific to the pronunciation of spoken words, our method adds critical dimensions to text transcripts. This integration also provides insights into the unique speech patterns of individuals, revealing the potential mental states of individuals. Evaluations of the proposed approach on the DAIC-WOZ dataset reveal state-of-the-art results when compared with existing Audio-Text baselines. In addition, this approach is not only valuable for the detection of depression but also represents a new perspective in enhancing the ability of LLMs to comprehend and process speech signals.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick 等ICLR 2022 · 被引用 1,182 次
相关 Paper
- DRMD: Explainable Depression Detection Based on Metaphorical Conceptual MappingDongyu Zhang, Wanqiu Liao, Weichen Hu, Hongfei LinWWW 2026
- DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's DiseaseTingyu Mo, Jacqueline C. K. Lam, Victor O. K. Li, Lawrence Y. L. CheungAAAI 2025 · 被引用 4 次
- Using Graph Representation Learning with Schema Encoders to Measure the Severity of Depressive SymptomsSimin Hong, Anthony G. Cohn, David Crossland HoggICLR 2022 · 被引用 16 次
- Mental-Perceiver: Audio-Textual Multi-Modal Learning for Estimating Mental DisordersJinghui Qin, Changsong Liu, Tianchi Tang, Dahuang Liu 等AAAI 2025 · 被引用 6 次
- Predicting Depression in Screening Interviews from Latent Categorization of Interview PromptsAlex Rinaldi, Jean E. Fox Tree, Snigdha ChaturvediACL 2020 · 被引用 20 次
