ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models
Haiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong, Jiaan Wang, Kexin Huang, Tianle Gu, Yixu Wang, Jian Wang, Dandan Liang, Zhixu Li, Yan Teng
摘要
Emotion Support Conversation (ESC) is a crucial application, which aims to reduce human stress, offer emotional guidance, and ultimately enhance human mental and physical well-being. With the advancement of Large Language Models (LLMs), many researchers have employed LLMs as the ESC models. However, the evaluation of these LLM-based ESCs remains uncertain. Inspired by the awesome development of role-playing agents, we propose an ESC Evaluation framework (i.e., ESC-Eval), which uses a role-playing agent to interact with ESC models, followed by a manual evaluation of the interactive dialogues. In detail, we first reorganize 2,801 role-playing cards from seven existing datasets to define the roles of the roleplaying agent. Second, we train a specific roleplaying model -ESC-Role to mimic the behavior of a real person experiencing distress. Third, through ESC-Role and organized role cards, we systematically conduct experiments using 14 LLMs as the ESC models, including general AI-assistant LLMs (e.g., ChatGPT) and ESC-oriented LLMs (e.g., ExTES-Llama). We conduct comprehensive human annotations on interactive multi-turn dialogues of different ESC models. The results show that ESCoriented LLMs exhibit superior ESC abilities compared to general AI-assistant LLMs, but there is still a gap behind human performance. Moreover, to automate the evaluation of future ESC models, we developed ESC-RANK, which trained on the annotated data, achieving a scoring performance surpassing 35 points of GPT-4. Our data and code are available at https://github.com/AIFlames/Esc-Eval .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- KokoroChat: A Japanese Psychological Counseling Dialogue Dataset Collected via Role-Playing by Trained CounselorsZhiyang Qi, Takumasa Kaneko, Keiko Takamizo, Mariko Ukiyo 等ACL 2025 · 被引用 7 次
- TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue AgentXingyu Sui, Yanyan Zhao, Yulin Hu, Jiahe Guo 等ACL 2026 · 被引用 3 次
- Breakdowns in Conversational AI: Interactional Failures in Emotionally and Ethically Sensitive ContextsJiawen Deng, Wentao Zhang, Ziyun Jiao, Fuji RenCHI 2026 · 被引用 3 次
- ESC-Judge: A Framework for Comparing Emotional Support Conversational AgentsNavid Madani, Rohini K. SrihariEMNLP 2025 · 被引用 2 次
- EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldJing Ye, Lu Xiang, Yaping Zhang, Chengqing ZongACL 2026 · 被引用 2 次
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- MISC: A Mixed Strategy-Aware Model integrating COMET for Emotional Support ConversationQuan Tu, Yanran Li, Jianwei Cui, Bin Wang 等ACL 2022 · 被引用 141 次
- Character-LLM: A Trainable Agent for Role-PlayingYunfan Shao, Linyang Li, Junqi Dai, Xipeng QiuEMNLP 2023 · 被引用 97 次
- A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health SupportAshish Sharma, Adam S. Miner, David C. Atkins, Tim AlthoffEMNLP 2020 · 被引用 21 次
- Towards Emotional Support Dialog SystemsSiyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour 等ACL 2021
相关 Paper
- Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon 等ACL 2024 · 被引用 14 次
- ESCA: An Emotional Support Conversation Agent for Enhancing Reasonable Strategy Planning and Effective ExpressionJing Li, Yanxin Luo, Donghong Han, Yimeng Zhan 等AAAI 2026
- ChatAnime: Towards User-Centered Emotional Support in LLM-based Virtual Character ChatLanlan Qiu, Sophia Xiao Pu, Yeqi Feng, Wenchang Gao 等ACL 2026
- Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support ConversationLin Zhong, Renjin Zhu, Shujuan Ma, Jinhao Cui 等ACL 2026
- Rethinking the Evaluation for Conversational Recommendation in the Era of Large Language ModelsXiaolei Wang, Xinyu Tang, Xin Zhao, Jingyuan Wang 等EMNLP 2023 · 被引用 69 次
