TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
Xingyu Sui, Yanyan Zhao, Yulin Hu, Jiahe Guo, Weixiang Zhao, Bing Qin
Abstract
Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance. However, existing ESC systems and benchmarks largely focus on affective support in text-only settings, overlooking how external tools can enable factual grounding and reduce hallucination in multi-turn emotional support. We introduce TEA-Bench, the first interactive benchmark for evaluating tool-augmented agents in ESC, featuring realistic emotional scenarios, an MCP-style tool environment, and process-level metrics that jointly assess the quality and factual grounding of emotional support. Experiments on nine LLMs show that tool augmentation generally improves emotional support quality and reduces hallucination, but the gains are strongly capacity-dependent: stronger models use tools more selectively and effectively, while weaker models benefit only marginally. We further release TEA-Dialog, a dataset of tool-enhanced ESC dialogues, and find that supervised fine-tuning improves indistribution support but generalizes poorly. Our results underscore the importance of tool use in building reliable emotional support agents. 1 * Corresponding author 1 Our code and data can be found in https://github. com/XingYuSSS/TEA-Bench .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5df009a6-719f-47b7-a27b-3075f40ee0b8Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- MISC: A Mixed Strategy-Aware Model integrating COMET for Emotional Support ConversationQuan Tu, Yanran Li, Jianwei Cui, Bin Wang et al.ACL 2022 · 141 citations
- API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMsMinghao Li, Yingxiu Zhao, Bowen Yu, Feifan Song et al.EMNLP 2023 · 72 citations
- Evaluating Very Long-Term Conversational Memory of LLM AgentsAdyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal et al.ACL 2024 · 30 citations
- ToolBeHonest: A Multi-level Hallucination Diagnostic Benchmark for Tool-Augmented Large Language ModelsYuxiang Zhang, Jing Chen, Junjie Wang, Yaxin Liu et al.EMNLP 2024 · 6 citations
Related papers
- ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional SupportTiantian Chen, Jiaqi Lu, Ying Shen, Lin ZhangWWW 2026 · 1 citation
- EmoHarbor: Evaluating Personalized Emotional Support by Simulating the User's Internal WorldJing Ye, Lu Xiang, Yaping Zhang, Chengqing ZongACL 2026 · 2 citations
- Towards Emotional Support Dialog SystemsSiyang Liu, Chujie Zheng, Orianna Demasi, Sahand Sabour et al.ACL 2021
- ESCA: An Emotional Support Conversation Agent for Enhancing Reasonable Strategy Planning and Effective ExpressionJing Li, Yanxin Luo, Donghong Han, Yimeng Zhan et al.AAAI 2026
- Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon et al.ACL 2024 · 14 citations
