Measuring What Matters!! Assessing Therapeutic Principles in Mental-Health Conversation
Abdullah Mazhar, Het Riteshkumar Shah, Aseem Srivastava, Smriti Joshi, Md. Shad Akhtar
Abstract
The increasing use of large language models in mental health applications calls for principled evaluation frameworks that assess alignment with psychotherapeutic best practices beyond surface-level fluency. While recent systems exhibit conversational competence, they lack structured mechanisms to evaluate adherence to core therapeutic principles. In this paper, we study the problem of evaluating AI-generated therapist-like responses for clinically grounded appropriateness and effectiveness. We assess each therapists utterance along six therapeutic principles: non-judgmental acceptance, warmth, respect for autonomy, active listening, reflective understanding, and situational appropriateness using a fine-grained ordinal scale. We introduce FAITH-M, a benchmark annotated with expert-assigned ordinal ratings, and propose CARE, a multi-stage evaluation framework that integrates intra-dialogue context, contrastive exemplar retrieval, and knowledge-distilled chain-of-thought reasoning. Experiments show that CARE achieves an F-1 score of 63.34 versus the strong baseline Qwen3 F-1 score of 38.56 which is a 64.26 improvement, which also serves as its backbone, indicating that gains arise from structured reasoning and contextual modeling rather than backbone capacity alone. Expert assessment and external dataset evaluations further demonstrate robustness under domain shift, while highlighting challenges in modelling implicit clinical nuance. Overall, CARE provides a clinically grounded framework for evaluating therapeutic fidelity in AI mental health systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c3f6c6c3-9f0d-49d6-a4f2-29f585f3bfecBuilds on8
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Towards Facilitating Empathic Conversations in Online Mental Health Support: A Reinforcement Learning ApproachAshish Sharma, Inna W. Lin, Adam S. Miner, David C. Atkins et al.WWW 2021 · 183 citations
- The Typing Cure: Experiences with Large Language Model Chatbots for Mental Health SupportInhwa Song, Sachin R. Pendse, Neha Kumar, Munmun De ChoudhuryCSCW 2025 · 41 citations
Related papers
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank et al.ICLR 2026 · 24 citations
- Design and Multi-level Evaluation of MAP-X: a Medically Aligned, Patient-Centered AI Explanation SystemYuyoung Kim, Minjung Kim, Saebyeol Kim, Sooyoun Cho et al.CHI 2026 · 1 citation
- Trustworthy Medical Question Answering: An Evaluation-Centric SurveyYinuo Wang, Baiyang Wang, Robert E. Mercer, Frank Rudzicz et al.EMNLP 2025 · 2 citations
- MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive ExplorationHao Lu, Yanchi Gu, Haoyuan Huang, Yulin Zhou et al.AAAI 2026 · 3 citations
- How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response DraftingParker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe et al.ACL 2026
