Towards Interpretable Mental Health Analysis with Large Language Models
Kailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie, Ziyan Kuang, Sophia Ananiadou
Abstract
The latest large language models (LLMs) such as ChatGPT, exhibit strong capabilities in automated mental health analysis. However, existing relevant studies bear several limitations, including inadequate evaluations, lack of prompting strategies, and ignorance of exploring LLMs for explainability. To bridge these gaps, we comprehensively evaluate the mental health analysis and emotional reasoning ability of LLMs on 11 datasets across 5 tasks. We explore the effects of different prompting strategies with unsupervised and distantly supervised emotional information. Based on these prompts, we explore LLMs for interpretable mental health analysis by instructing them to generate explanations for each of their decisions. We convey strict human evaluations to assess the quality of the generated explanations, leading to a novel dataset with 163 humanassessed explanations 1 . We benchmark existing automatic evaluation metrics on this dataset to guide future related works. According to the results, ChatGPT shows strong in-context learning ability but still has a significant gap with advanced task-specific methods. Careful prompt engineering with emotional cues and expertwritten few-shot examples can also effectively improve performance on mental health analysis. In addition, ChatGPT generates explanations that approach human performance, showing its great potential in explainable mental health analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a6237054-932c-4b10-ad6b-288054d2d3baCited by top-tier papers17
- MetaAligner: Towards Generalizable Multi-Objective Alignment of Language ModelsKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang et al.NeurIPS 2024 · 49 citations
- "Kya family planning after marriage hoti hai?": Integrating Cultural Sensitivity in an LLM Chatbot for Reproductive HealthRoshini Deva, Dhruv Ramani, Tanvi Divate, Suhani Jalota et al.CHI 2025 · 16 citations
- Exploring Modular Prompt Design for Emotion and Mental Health RecognitionMinseo Kim, Taemin Kim, Thu Hoang Anh Vo, Yugyeong Jung et al.CHI 2025 · 11 citations
- Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health ConversationsMohit Chandra, Siddharth Sriraman, Harneet Singh Khanuja, Yiqiao Jin et al.ACL 2026 · 6 citations
- MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with ResistanceSubin Kim, Hoonrae Kim, Jihyun Lee, Yejin Jeon et al.EMNLP 2025 · 6 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social MediaWei Zhai, Nan Bai, Qing Zhao, Jianqiang Li et al.EMNLP 2025
- Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text DataXuhai Xu, Bingsheng Yao, Yuanzhe Dong, Saadia Gabriel et al.UbiComp 2024 · 281 citations
- A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsChen Zhang, Luis Fernando D'Haro, Yiming Chen, Malu Zhang et al.AAAI 2024 · 57 citations
- 'I've talked to ChatGPT about my issues last night.': Examining Mental Health Conversations with Large Language Models through Reddit AnalysisKyuha Jung, Gyuho Lee, Yuanhui Huang, Yunan ChenCSCW 2025 · 23 citations
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
