Responsible Evaluation of AI for Mental Health
Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah
Abstract
Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user experience. This paper argues for a rethinking of responsible evaluation -- what is measured, by whom, and for what purpose -- by introducing an interdisciplinary framework that integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. Through an analysis of 135 recent *CL publications, we identify recurring limitations, including over-reliance on generic metrics that do not capture clinical validity, therapeutic appropriateness, or user experience, limited participation from mental health professionals, and insufficient attention to safety and equity. To address these gaps, we propose a taxonomy of AI mental health support types -- assessment-, intervention-, and information synthesis-oriented -- each with distinct risks and evaluative requirements, and illustrate its use through case studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05ece280-33bb-4b9b-8a5b-eaf80fe13fddBuilds on40
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie et al.EMNLP 2023 · 114 citations
- Improving the Generalizability of Depression Detection by Leveraging Clinical QuestionnairesThong Nguyen, Andrew Yates, Ayah Zirikly, Bart Desmet et al.ACL 2022 · 66 citations
- Identifying Moments of Change from Longitudinal User TextAdam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim et al.ACL 2022 · 46 citations
- PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological CounselingHaojie Xie, Yirong Chen, Xiaofen Xing, Jingkai Lin et al.ACL 2025 · 43 citations
- Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to PrinciplesRyan Louie, Ananjan Nandi, William Fang, Cheng Chang et al.EMNLP 2024 · 37 citations
Related papers
- Framing Responsible Design of AI for Mental Well-Being: AI as Primary Care, Nutritional Supplement, or Yoga Instructor?Ned Cooper, Jose A. Guridi, Angel Hsing-Chi Hwang, Beth Kolko et al.CHI 2026 · 1 citation
- Cloning the Self for Mental Well-Being: A Framework for Designing Safe and Therapeutic Self-Clone ChatbotsMehrnoosh Sadat Shirvani, Jackie Crowley, Cher Peng, Jackie Liu et al.CHI 2026 · 1 citation
- A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness EvaluationsGlen Berman, Nitesh Goyal, Michael MadaioCHI 2024 · 40 citations
- From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating DisordersAmy A. Winecoff, Kevin KlymanCHI 2026
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health SupportInhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De ChoudhuryCSCW 2025 · 41 citations
