An In-depth Investigation of User Response Simulation for Conversational Search
Zhenduo Wang, Zhichao Xu, Vivek Srikumar, Qingyao Ai
Abstract
Conversational search has seen increased recent attention in both the IR and NLP communities. It seeks to clarify and solve users' search needs through multi-turn natural language interactions. However, most existing systems are trained and demonstrated with recorded or artificial conversation logs. Eventually, conversational search systems should be trained, evaluated, and deployed in an open-ended setting with unseen conversation trajectories. A key challenge is that training and evaluating such systems both require a human-in-the-loop, which is expensive and does not scale. One strategy is to simulate users, thereby reducing the scaling costs. However, current user simulators are either limited to only responding to yes-no questions from the conversational search system or unable to produce high-quality responses in general. This paper shows that existing user simulation systems could be significantly improved by a smaller finetuned natural language generation model. However, rather than merely reporting it as the new state-of-the-art, we consider it a strong baseline and present an in-depth investigation of simulating user response for conversational search. Our goal is to supplement existing work with an insightful hand-analysis of unsolved challenges by the baseline and propose our solutions. The challenges we identified include (1) a blind spot that is difficult to learn, and (2) a specific type of misevaluation in the standard setup. We propose a new generation system to effectively cover the training blind spot and suggest a new evaluation setup to avoid misevaluation. Our proposed system leads to significant improvements over existing systems and large language models such as GPT-4. Additionally, our analysis provides insights into the nature of the task to facilitate future work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 444ad3af-4099-48e5-b729-fdba97b117a4Cited by top-tier papers3
- MemoryBench: A Benchmark for Memory and Continual Learning in LLM SystemsQingyao Ai, Yichen Tang, Changyue Wang, Jianming Long et al.ICML 2026 · 47 citations
- Personality-aware Student Simulation for Conversational Intelligent Tutoring SystemsZhengyuan Liu, Stella Xin Yin, Geyu Lin, Nancy F. ChenEMNLP 2024 · 13 citations
- Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in RecommendationTianjun Wei, Huizhong Guo, Yingpeng Du, Zhu Sun et al.ACL 2026 · 4 citations
Builds on7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Generating Clarifying Questions for Information RetrievalHamed Zamani, Susan T. Dumais, Nick Craswell, Paul N. Bennett et al.WWW 2020 · 238 citations
- Picking on the Same Person: Does Algorithmic Monoculture lead to Outcome Homogenization?Rishi Bommasani, Kathleen A. Creel, Ananya Kumar, Dan Jurafsky et al.NeurIPS 2022 · 179 citations
- Analyzing and Learning from User Interactions for Search ClarificationHamed Zamani, Bhaskar Mitra, Everest Chen, Gord Lueck et al.SIGIR 2020 · 84 citations
Related papers
- Exploiting Simulated User Feedback for Conversational Search: Ranking, Rewriting, and BeyondPaul Owoicho, Ivan Sekulic, Mohammad Aliannejadi, Jeffrey Dalton et al.SIGIR 2023 · 31 citations
- Do Simulated Users Need to Remember? Analyzing the Impact of Memory Models in Conversational Search EvaluationNailia Mirzakhmedova, Marcel Gohsen, Johannes Kiesel, Matthias Hagen et al.SIGIR 2026
- Evaluating Conversational Recommender Systems via User SimulationShuo Zhang, Krisztian BalogKDD 2020 · 80 citations
- UniConv: Unifying Retrieval and Response Generation for Large Language Models in ConversationsFengran Mo, Yifan Gao, Chuan Meng, Xin Liu et al.ACL 2025 · 22 citations
- Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking AgentsWeiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang et al.EMNLP 2023 · 182 citations
