Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis
Amey Hengle, Atharva Kulkarni, Shantanu Patankar, Madhumitha Chandrasekaran, Sneha D'Silva, Jemima Jacob, Rashmi Gupta
摘要
Warning: This paper includes examples displaying symptoms of mental health disorders for contextual understanding. In this study, we introduce ANGST, a novel, first of its kind benchmark for depressionanxiety comorbidity classification from social media posts. Unlike contemporary datasets that often oversimplify the intricate interplay between different mental health disorders by treating them as isolated conditions, ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety. Comprising 2876 meticulously annotated posts by expert psychologists and an additional 7667 silver-labeled posts, ANGST posits a more representative sample of online mental health discourse. Moreover, we benchmark ANGST using various state-of-the-art language models, ranging from Mental-BERT to GPT-4. Our results provide significant insights into the capabilities and limitations of these models in complex diagnostic scenarios. While GPT-4 generally outperforms other models, none achieve an F1 score exceeding 72% in multiclass comorbid classification, underscoring the ongoing challenges in applying language models to mental health diagnostics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- A Multi-Level Benchmark for Causal Language Understanding in Social Media DiscourseXiaohan Ding, Kaike Ping, Buse Çarik, Eugenia Ha Rim RhoEMNLP 2025 · 被引用 1 次
- MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric DiagnosisXiao Sun, Yuming Yang, Xinyi Jiang, Yu Tian 等ACL 2026 · 被引用 1 次
- Responsible Evaluation of AI for Mental HealthHiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt 等ACL 2026
- ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social MediaAakash Kumar Agarwal, Saprativa Bhattacharjee, Mauli Rastogi, Jemima Jacob 等EMNLP 2025
它引用的顶会 Paper5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel 等UbiComp 2024 · 被引用 281 次
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie 等EMNLP 2023 · 被引用 114 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Improving the Generalizability of Depression Detection by Leveraging Clinical QuestionnairesThong Nguyen, Andrew Yates, Ayah Zirikly, Bart Desmet 等ACL 2022 · 被引用 66 次
相关 Paper
- Exploring Large Language Models for Detecting Mental DisordersGleb Kuzmin, Petr Strepetov, Maksim Stankevich, Natalya V. Chudova 等EMNLP 2025
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language ModelsZachary Englhardt, Chengqian Ma, Margaret E. Morris, Chun-Cheng Chang 等UbiComp 2024 · 被引用 51 次
- Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme ClassificationAbdullah Mazhar, Zuhair Hasan Shaik, Aseem Srivastava, Polly Ruhnke 等WWW 2025 · 被引用 9 次
- DRMD: Explainable Depression Detection Based on Metaphorical Conceptual MappingDongyu Zhang, Wanqiu Liao, Weichen Hu, Hongfei LinWWW 2026
- Symptom Detection with Text Message Log Distributions for Holistic Depression and Anxiety ScreeningM. L. Tlachac, Michael V. Heinz, Miranda Reisch, Samuel S. OgdenUbiComp 2024 · 被引用 5 次
