ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents
Navid Madani, Rohini K. Srihari
摘要
Large Language Models (LLMs) increasingly power mental-health chatbots, yet the field still lacks a scalable, theory-grounded way to decide which model is more effective to deploy. We present ESC-Judge, the first end-to-end evaluation framework that (i) grounds head-tohead comparison of Emotional-Support LLMs (ES-LLMs) in an established psychological theory-Clara Hill's Exploration-Insight-Action (E-I-A) counselling model-thereby delivering a structured, interpretable lens on performance, and (ii) fully automates the pipeline at scale. ESC-Judge proceeds in three stages: (1) it synthesizes realistic help-seeker roles by sampling empirically salient attributes (stressors, personality, life history); (2) it has two candidate ES-Agents conduct separate sessions with the same role, isolating model-specific strategies; and (3) it asks a specialised judge LLM to issue pairwise preferences across rubric-anchored skills that exhaustively cover the E-I-A spectrum. In our empirical study, ESC-Judge matches PhDlevel annotators in 85% of Exploration, 83% of Insight, and 86% of Action decisions, demonstrating human-level reliability at a fraction of the cost. We release all code, prompts, synthetic roles, transcripts, and judgment scripts to catalyze transparent progress in emotionally supportive AI 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive ContextsJiawen Deng, Wentao Zhang, Ziyun Jiao, Fuji RenCHI 2026 · 被引用 3 次
- EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldJing Ye, Lu Xiang, Yaping Zhang, Chengqing ZongACL 2026 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite FeedbackShijing Zhu, Zhuang Chen, Guanqun Bi, Binghang Li 等AAAI 2026
- An LLM-based Simulation Framework for Embodied Conversational Agents in Psychological CounselingLixiu Wu, Yuanrong Tang, Qisen Pan, Xianyang Zhan 等AAAI 2026 · 被引用 1 次
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie 等EMNLP 2023 · 被引用 114 次
- Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon 等ACL 2024 · 被引用 14 次
- SocialSim: Towards Socialized Simulation of Emotional Support ConversationZhuang Chen, Yaru Cao, Guanqun Bi, Jincenzi Wu 等AAAI 2025 · 被引用 12 次
