Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language Models
Zhouxing Tan, Ruochong Xiong, Yulong Wan, Jinlong Ma, Hanlin Xue, Qichun Deng, Haifeng Jing, Zhengtong Zhang, Depei Liu, Shiyuan Luo, Junfei Liu
摘要
Emotional support is a core capability in human-AI interaction, with applications including psychological counseling, role play, and companionship. However, existing evaluations of large language models (LLMs) often rely on short, static dialogues and fail to capture the dynamic and long-term nature of emotional support. To overcome this limitation, we shift from snapshot-based evaluation to trajectory-based assessment, adopting a user-centered perspective that evaluates models based on their ability to improve and stabilize user emotional states over time. Our framework constructs a large-scale benchmark consisting of 328 emotional contexts and 1,152 disturbance events, simulating realistic emotional shifts under evolving dialogue scenarios. To encourage psychologically grounded responses, we constrain model outputs using validated emotion regulation strategies such as situation selection and cognitive reappraisal. User emotional trajectories are modeled as a first-order Markov process, and we apply causally-adjusted emotion estimation to obtain unbiased emotional state tracking. Based on this framework, we introduce three trajectory-level metrics: Baseline Emotional Level (BEL), Emotional Trajectory Volatility (ETV), and Emotional Centroid Position (ECP). These metrics collectively capture user emotional dynamics over time and support comprehensive evaluation of long-term emotional support performance of LLMs. Extensive evaluations across a diverse set of LLMs reveal significant disparities in emotional support capabilities and provide actionable insights for model development.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI SynergyChris Yuhao Liu, Liang Zeng, Yuzhen Xiao, Jujie He 等ICLR 2026 · 被引用 211 次
- Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon 等ACL 2024 · 被引用 14 次
- CharacterBench: Benchmarking Character Customization of Large Language ModelsJinfeng Zhou, Yongkang Huang, Bosi Wen, Guanqun Bi 等AAAI 2025 · 被引用 7 次
- DMT-RoleBench: A Dynamic Multi-Turn Dialogue Based Benchmark for Role-Playing Evaluation of Large Language Model and AgentDingbo Yuan, Yipeng Chen, Guodong Liu, Chenchen Li 等AAAI 2025 · 被引用 6 次
相关 Paper
- Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human StatesYang Xiao, Jiashuo Wang, Qiancheng Xu, Changhe Song 等ACL 2025 · 被引用 12 次
- ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional SupportTiantian Chen, Jiaqi Lu, Ying Shen, Lin ZhangWWW 2026 · 被引用 1 次
- From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support ConversationsShenghan Wu, Yimo Zhu, Wynne Hsu, Mong-Li Lee 等EMNLP 2025 · 被引用 1 次
- ESC-Eval: Evaluating Emotion Support Conversations in Large Language ModelsHaiquan Zhao, Lingyu Li, Shisong Chen, Shuqi Kong 等EMNLP 2024 · 被引用 5 次
- EmoBench: Evaluating the Emotional Intelligence of Large Language ModelsSahand Sabour, Siyang Liu, Zheyuan Zhang, June M. Liu 等ACL 2024
