Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis
Amey Hengle, Atharva Kulkarni, Shantanu Patankar, Madhumitha Chandrasekaran, Sneha D'Silva, Jemima Jacob, Rashmi Gupta
Abstract
Warning: This paper includes examples displaying symptoms of mental health disorders for contextual understanding. In this study, we introduce ANGST, a novel, first of its kind benchmark for depressionanxiety comorbidity classification from social media posts. Unlike contemporary datasets that often oversimplify the intricate interplay between different mental health disorders by treating them as isolated conditions, ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety. Comprising 2876 meticulously annotated posts by expert psychologists and an additional 7667 silver-labeled posts, ANGST posits a more representative sample of online mental health discourse. Moreover, we benchmark ANGST using various state-of-the-art language models, ranging from Mental-BERT to GPT-4. Our results provide significant insights into the capabilities and limitations of these models in complex diagnostic scenarios. While GPT-4 generally outperforms other models, none achieve an F1 score exceeding 72% in multiclass comorbid classification, underscoring the ongoing challenges in applying language models to mental health diagnostics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e4144bad-201a-427b-acc6-90c19bdbf66bCited by top-tier papers4
- A Multi-Level Benchmark for Causal Language Understanding in Social Media DiscourseXiaohan Ding, Kaike Ping, Buse Çarik, Eugenia Ha Rim RhoEMNLP 2025 · 1 citation
- MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric DiagnosisXiao Sun, Yuming Yang, Xinyi Jiang, Yu Tian et al.ACL 2026 · 1 citation
- Responsible Evaluation of AI for Mental HealthHiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt et al.ACL 2026
- ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social MediaAakash Kumar Agarwal, Saprativa Bhattacharjee, Mauli Rastogi, Jemima Jacob et al.EMNLP 2025
Builds on5
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel et al.UbiComp 2024 · 281 citations
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie et al.EMNLP 2023 · 114 citations
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo et al.ACL 2020 · 93 citations
- Improving the Generalizability of Depression Detection by Leveraging Clinical QuestionnairesThong Nguyen, Andrew Yates, Ayah Zirikly, Bart Desmet et al.ACL 2022 · 66 citations
Related papers
- Exploring Large Language Models for Detecting Mental DisordersGleb Kuzmin, Petr Strepetov, Maksim Stankevich, Natalya V. Chudova et al.EMNLP 2025
- From Classification to Clinical Insights: Towards Analyzing and Reasoning About Mobile and Behavioral Health Data With Large Language ModelsZachary Englhardt, Chengqian Ma, Margaret E. Morris, Chun-Cheng Chang et al.UbiComp 2024 · 51 citations
- Figurative-cum-Commonsense Knowledge Infusion for Multimodal Mental Health Meme ClassificationAbdullah Mazhar, Zuhair Hasan Shaik, Aseem Srivastava, Polly Ruhnke et al.WWW 2025 · 9 citations
- DRMD: Explainable Depression Detection Based on Metaphorical Conceptual MappingDongyu Zhang, Wanqiu Liao, Weichen Hu, Hongfei LinWWW 2026
- Symptom Detection with Text Message Log Distributions for Holistic Depression and Anxiety ScreeningM. L. Tlachac, Michael V. Heinz, Miranda Reisch, Samuel S. OgdenUbiComp 2024 · 5 citations
