EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal World
Jing Ye, Lu Xiang, Yaping Zhang, Chengqing Zong
Abstract
Current evaluation paradigms for emotional support conversations tend to reward generic empathetic responses, yet they fail to assess whether the support is genuinely personalized to users' unique psychological profiles and contextual needs. We introduce EmoHarbor, an automated evaluation framework that adopts a User-as-a-Judge paradigm by simulating the user's inner world. EmoHarbor employs a Chain-of-Agent architecture that decomposes users' internal processes into three specialized roles, enabling agents to interact with supporters and complete assessments in a manner similar to human users. We instantiate this benchmark using 100 real-world user profiles that cover a diverse range of personality traits and situations, and define 10 evaluation dimensions of personalized support quality. Comprehensive evaluation of 20 advanced LLMs on EmoHarbor reveals a critical insight: while these models excel at generating empathetic responses, they consistently fail to tailor support to individual user contexts. This finding reframes the central challenge, shifting research focus from merely enhancing generic empathy to developing truly user-aware emotional support. EmoHarbor provides a reproducible and scalable framework to guide the development and evaluation of more nuanced and user-aware emotional support systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb3861d3-6506-4f86-a95c-62ba5fc230eeBuilds on14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang et al.ICLR 2024 · 288 citations
- Customizing Emotional Support: How Do Individuals Construct and Interact With LLM-Powered ChatbotsXi Zheng, Zhuoyang Li, Xinning Gui, Yuhan LuoCHI 2025 · 47 citations
Related papers
- ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional SupportTiantian Chen, Jiaqi Lu, Ying Shen, Lin ZhangWWW 2026 · 1 citation
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language ModelsLi Zhou, Lutong Yu, You Lyu, Yihang Lin et al.ICLR 2026 · 13 citations
- Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language ModelsZhouxing Tan, Ruochong Xiong, Yulong Wan, Jinlong Ma et al.AAAI 2026
- TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue AgentXingyu Sui, Yanyan Zhao, Yulin Hu, Jiahe Guo et al.ACL 2026 · 3 citations
- Can LLM Agents Maintain a Persona in Discourse?Pranav Bhandari, Nicolas Fay, Michael J. Wise, Amitava Datta et al.EMNLP 2025
