ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents
Navid Madani, Rohini K. Srihari
Abstract
Large Language Models (LLMs) increasingly power mental-health chatbots, yet the field still lacks a scalable, theory-grounded way to decide which model is more effective to deploy. We present ESC-Judge, the first end-to-end evaluation framework that (i) grounds head-tohead comparison of Emotional-Support LLMs (ES-LLMs) in an established psychological theory-Clara Hill's Exploration-Insight-Action (E-I-A) counselling model-thereby delivering a structured, interpretable lens on performance, and (ii) fully automates the pipeline at scale. ESC-Judge proceeds in three stages: (1) it synthesizes realistic help-seeker roles by sampling empirically salient attributes (stressors, personality, life history); (2) it has two candidate ES-Agents conduct separate sessions with the same role, isolating model-specific strategies; and (3) it asks a specialised judge LLM to issue pairwise preferences across rubric-anchored skills that exhaustively cover the E-I-A spectrum. In our empirical study, ESC-Judge matches PhDlevel annotators in 85% of Exploration, 83% of Insight, and 86% of Action decisions, demonstrating human-level reliability at a fraction of the cost. We release all code, prompts, synthetic roles, transcripts, and judgment scripts to catalyze transparent progress in emotionally supportive AI 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive ContextsJiawen Deng, Wentao Zhang, Ziyun Jiao, Fuji RenCHI 2026 · 3 citations
- EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldJing Ye, Lu Xiang, Yaping Zhang, Chengqing ZongACL 2026 · 2 citations
Builds on2
Related papers
- Ψ-Arena: Interactive Assessment and Optimization of LLM-based Psychological Counselors with Tripartite FeedbackShijing Zhu, Zhuang Chen, Guanqun Bi, Binghang Li et al.AAAI 2026
- An LLM-based Simulation Framework for Embodied Conversational Agents in Psychological CounselingLixiu Wu, Yuanrong Tang, Qisen Pan, Xianyang Zhan et al.AAAI 2026 · 1 citation
- Towards Interpretable Mental Health Analysis with Large Language ModelsKailai Yang, Shaoxiong Ji, Tianlin Zhang, Qianqian Xie et al.EMNLP 2023 · 114 citations
- Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon et al.ACL 2024 · 14 citations
- SocialSim: Towards Socialized Simulation of Emotional Support ConversationZhuang Chen, Yaru Cao, Guanqun Bi, Jincenzi Wu et al.AAAI 2025 · 12 citations
