Do Simulated Users Need to Remember? Analyzing the Impact of Memory Models in Conversational Search Evaluation
Nailia Mirzakhmedova, Marcel Gohsen, Johannes Kiesel, Matthias Hagen, Benno Stein
Abstract
Conversational search systems are typically evaluated using a fixed reference collection of conversations or through user studies with a live system. However, fixed-reference conversations can cover only a few plausible conversations, and user studies are costly, time-consuming, and often hard to reproduce. A promising alternative that avoids coverage and cost issues is user simulation, in which a computer program takes on the role of a user and interacts with the system under evaluation. But the complexity of human search behavior raises the question of how ''realistic'' the simulations actually need to be for reliable evaluations of conversational search systems. In this paper, we ask: Do simulated users need to remember? While real users may learn and forget information during conversational search sessions, which inspired previous research to also model memory capabilities in simulations, it remains unclear whether this actually influences the results of system evaluations. To investigate the impact of memory modeling, we analyze conversations of simulated users and of humans with four conversational search systems. Our results suggest that incorporating long-term memory into simulators can help reproduce system effectiveness rankings obtained from human conversations, whereas incorporating short-term memory can diminish the reproduction. We also find that simulators are generally valid and reproducible---and memory modeling even increases run-to-run reproducibility of system rankings---but overall, simulations approximate human evaluation scores better when ''helpful'' assistants are evaluated than when assistants with deteriorated response quality are assessed. Our code and data are available at https://github.com/webis-de/SIGIR-26.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ecc830bf-2cd8-43cc-b40f-ac97e507ac53Builds on7
- MemoryBank: Enhancing Large Language Models with Long-Term MemoryWanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye et al.AAAI 2024 · 394 citations
- Evaluating Conversational Recommender Systems via User SimulationShuo Zhang, Krisztian BalogKDD 2020 · 80 citations
- Exploiting Simulated User Feedback for Conversational Search: Ranking, Rewriting, and BeyondPaul Owoicho, Ivan Sekulic, Mohammad Aliannejadi, Jeffrey Dalton et al.SIGIR 2023 · 31 citations
- Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesKuang Wang, Xianfei Li, Shenghao Yang, Li Zhou et al.ACL 2025 · 24 citations
- A LLM-based Controllable, Scalable, Human-Involved User Simulator Framework for Conversational Recommender SystemsLixi Zhu, Xiaowen Huang, Jitao SangWWW 2025 · 18 citations
Related papers
- An In-depth Investigation of User Response Simulation for Conversational SearchZhenduo Wang, Zhichao Xu, Vivek Srikumar, Qingyao AiWWW 2024 · 32 citations
- Understanding How Psychological Distance Influences User Preferences in Conversational versus Web SearchYitian Yang, Yugin Tan, Yang Chen Lin, Jung-Tai King et al.CHI 2025 · 11 citations
- Task-Aware Automated User Profile Generation for Recommendation Simulation Using Large Language ModelsXinye Wanyan, Chenglong Ma, Danula Hettiachchi, Ziqi Xu et al.SIGIR 2026
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive MemoryDi Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang et al.ICLR 2025
- Simulating Human Imprecision in Temporal Statements of Intelligent Virtual AgentsSusanne Schmidt, Sven Zimmermann, Celeste Mason, Frank SteinickeCHI 2022 · 6 citations
