ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
Tiantian Chen, Jiaqi Lu, Ying Shen, Lin Zhang
摘要
Large Language Models (LLMs) have shown strong potential as conversational agents. Yet, their effectiveness remains limited by deficiencies in robust long-term memory—particularly in complex, long-term Web-based services such as online emotional support. However, existing long-term dialogue benchmarks primarily focus on static and explicit fact retrieval, failing to evaluate agents in these critical scenarios where user information is dispersed, implicit, and continuously evolving. To address this gap, we introduce ES-MemEval, a comprehensive benchmark that systematically evaluates five core memory capabilities—information extraction, temporal reasoning, conflict detection, abstention, and user modeling—in long-term emotional support scenarios, covering question answering, summarization, and dialogue generation tasks. To support the benchmark, we also propose EvoEmo, the first multi-session dataset for personalized long-term emotional support scenarios, capturing fragmented, implicit user disclosures and evolving user states. Extensive experiments on open-source long-context, commercial, and retrieval-augmented (RAG) LLMs reveal that explicit long-term memory is essential to reduce hallucinations and enable effective personalization. At the same time, RAG enhances factual consistency but struggles with temporal dynamics and evolving user states. These findings highlight both the potential and limitations of current paradigms, encouraging the development of more robust memory–retrieval integration in long-term personalized dialogue systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
相关 Paper
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryDi Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang 等ICLR 2025
- Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM AgentsYuanchen Bei, Tianxin Wei, Xuying Ning, Yanjun Zhao 等ACL 2026 · 被引用 23 次
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 被引用 246 次
- In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue AgentsZhen Tan, Jun Yan, I-Hung Hsu, Rujun Han 等ACL 2025
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMsSiyan Zhao, Mingyi Hong, Yang Liu, Devamanyu Hazarika 等ICLR 2025
