Responsible Evaluation of AI for Mental Health
Hiba Arnaout, Anmol Goel, H. Andrew Schwartz, Steffen Eberhardt, Dana Atzil-Slonim, Gavin Doherty, Brian Schwartz, Wolfgang Lutz, Tim Althoff, Munmun De Choudhury, Hamidreza Jamalabadi, Raj Sanjay Shah
摘要
Although artificial intelligence (AI) shows growing promise for mental health care, current approaches to evaluating AI tools in this domain remain fragmented and poorly aligned with clinical practice, social context, and first-hand user experience. This paper argues for a rethinking of responsible evaluation -- what is measured, by whom, and for what purpose -- by introducing an interdisciplinary framework that integrates clinical soundness, social context, and equity, providing a structured basis for evaluation. Through an analysis of 135 recent *CL publications, we identify recurring limitations, including over-reliance on generic metrics that do not capture clinical validity, therapeutic appropriateness, or user experience, limited participation from mental health professionals, and insufficient attention to safety and equity. To address these gaps, we propose a taxonomy of AI mental health support types -- assessment-, intervention-, and information synthesis-oriented -- each with distinct risks and evaluative requirements, and illustrate its use through case studies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper40
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie 等EMNLP 2023 · 被引用 114 次
- Improving the Generalizability of Depression Detection by Leveraging Clinical QuestionnairesThong Nguyen, Andrew Yates, Ayah Zirikly, Bart Desmet 等ACL 2022 · 被引用 66 次
- Identifying Moments of Change from Longitudinal User TextAdam Tsakalidis, Federico Nanni, Anthony Hills, Jenny Chim 等ACL 2022 · 被引用 46 次
- PsyDT: Using LLMs to Construct the Digital Twin of Psychological Counselor with Personalized Counseling Style for Psychological CounselingHaojie Xie, Yirong Chen, Xiaofen Xing, Jingkai Lin 等ACL 2025 · 被引用 43 次
- Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to PrinciplesRyan Louie, Ananjan Nandi, William Fang, Cheng Chang 等EMNLP 2024 · 被引用 37 次
相关 Paper
- Framing Responsible Design of AI for Mental Well-Being: AI as Primary Care, Nutritional Supplement, or Yoga Instructor?Ned Cooper, Jose A. Guridi, Angel Hsing-Chi Hwang, Beth Kolko 等CHI 2026 · 被引用 1 次
- Cloning the Self for Mental Well-Being: A Framework for Designing Safe and Therapeutic Self-Clone ChatbotsMehrnoosh Sadat Shirvani, Jackie Crowley, Cher Peng, Jackie Liu 等CHI 2026 · 被引用 1 次
- A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness EvaluationsGlen Berman, Nitesh Goyal, Michael MadaioCHI 2024 · 被引用 40 次
- From Symptoms to Systems: An Expert-Guided Approach to Understanding Risks of Generative AI for Eating DisordersAmy A. Winecoff, Kevin KlymanCHI 2026
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health SupportInhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De ChoudhuryCSCW 2025 · 被引用 41 次
