Do Simulated Users Need to Remember? Analyzing the Impact of Memory Models in Conversational Search Evaluation
Nailia Mirzakhmedova, Marcel Gohsen, Johannes Kiesel, Matthias Hagen, Benno Stein
摘要
Conversational search systems are typically evaluated using a fixed reference collection of conversations or through user studies with a live system. However, fixed-reference conversations can cover only a few plausible conversations, and user studies are costly, time-consuming, and often hard to reproduce. A promising alternative that avoids coverage and cost issues is user simulation, in which a computer program takes on the role of a user and interacts with the system under evaluation. But the complexity of human search behavior raises the question of how ''realistic'' the simulations actually need to be for reliable evaluations of conversational search systems. In this paper, we ask: Do simulated users need to remember? While real users may learn and forget information during conversational search sessions, which inspired previous research to also model memory capabilities in simulations, it remains unclear whether this actually influences the results of system evaluations. To investigate the impact of memory modeling, we analyze conversations of simulated users and of humans with four conversational search systems. Our results suggest that incorporating long-term memory into simulators can help reproduce system effectiveness rankings obtained from human conversations, whereas incorporating short-term memory can diminish the reproduction. We also find that simulators are generally valid and reproducible---and memory modeling even increases run-to-run reproducibility of system rankings---but overall, simulations approximate human evaluation scores better when ''helpful'' assistants are evaluated than when assistants with deteriorated response quality are assessed. Our code and data are available at https://github.com/webis-de/SIGIR-26.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye 等AAAI 2024 · 被引用 394 次
- Evaluating Conversational Recommender Systems via User SimulationShuo Zhang, Krisztian BalogKDD 2020 · 被引用 80 次
- Exploiting Simulated User Feedback for Conversational Search: Ranking, Rewriting, and BeyondPaul Owoicho, Ivan Sekulic, Mohammad Aliannejadi, Jeffrey Dalton 等SIGIR 2023 · 被引用 31 次
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesKuang Wang, Xianfei Li, Shenghao Yang, Li Zhou 等ACL 2025 · 被引用 24 次
- A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender SystemsLixi Zhu, Xiaowen Huang, Jitao SangWWW 2025 · 被引用 18 次
相关 Paper
- An In-depth Investigation of User Response Simulation for Conversational SearchZhenduo Wang, Zhichao Xu, Vivek Srikumar, Qingyao AiWWW 2024 · 被引用 32 次
- Understanding How Psychological Distance Influences User Preferences in Conversational versus Web SearchYitian Yang, Yugin Tan, Yang Chen Lin, Jung-Tai King 等CHI 2025 · 被引用 11 次
- Task-Aware Automated User Profile Generation for Recommendation Simulation Using Large Language ModelsXinye Wanyan, Chenglong Ma, Danula Hettiachchi, Ziqi Xu 等SIGIR 2026
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryDi Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang 等ICLR 2025
- Simulating Human Imprecision in Temporal Statements of Intelligent Virtual AgentsSusanne Schmidt, Sven Zimmermann, Celeste Mason, Frank SteinickeCHI 2022 · 被引用 6 次
