LLMs Get Lost In Multi-Turn Conversation
Philippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer Neville
Abstract
Large Language Models (LLMs) are conversational interfaces. As such, LLMs have the potential to assist their users not only when they can fully specify the task at hand, but also to help them define, explore, and refine what they need through multi-turn conversational exchange. Although analysis of LLM conversation logs has confirmed that underspecification occurs frequently in user instructions, LLM evaluation has predominantly focused on the single-turn, fully-specified instruction setting. In this work, we perform large-scale simulation experiments to compare LLM performance in single- and multi-turn settings. Our experiments confirm that all the top open- and closed-weight LLMs we test exhibit significantly lower performance in multi-turn conversations than single-turn, with an average drop of 39% across six generation tasks. Analysis of 200,000+ simulated conversations decomposes the performance degradation into two components: a minor loss in aptitude and a significant increase in unreliability. We find that LLMs often make assumptions in early turns and prematurely attempt to generate final solutions, on which they overly rely. In simpler terms, we discover that when LLMs take a wrong turn in a conversation, they get lost and do not recover.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd98a71b-c2ca-4d43-8352-b668d2f84c03Cited by top-tier papers59
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMsMohammad Tavakoli, Alireza Salemi, Carrie Ye, Mohamed Abdalla et al.ICLR 2026 · 56 citations
- Flipping the Dialogue: Training and Evaluating User Language ModelsTarek Naous, Philippe Laban, Wei Xu, Jennifer NevilleICLR 2026 · 56 citations
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic TasksShuo He, Lang Feng, Qi Wei, Xin Cheng et al.ICLR 2026 · 36 citations
- BED-LLM: Intelligent Information Gathering with LLMs and Bayesian Experimental DesignDeepro Choudhury, Sinead Williamson, Adam Golinski, Ning Miao et al.ICLR 2026 · 24 citations
- StreamingThinker: Large Language Models Can Think While ReadingJunlong Tong, Yingqi Fan, Anhao Zhao, Yunpu Ma et al.ICLR 2026 · 17 citations
Builds on23
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 892 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language FeedbackXingyao Wang, Zihan Wang, Jiateng Liu, Yangyi Chen et al.ICLR 2024 · 308 citations
- Design Principles for Generative AI ApplicationsJustin D. Weisz, Jessica He, Michael J. Muller, Gabriela Hoefer et al.CHI 2024 · 221 citations
Related papers
- MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language ModelsWai-Chung Kwan, Xingshan Zeng, Yuxin Jiang, Yufei Wang et al.EMNLP 2024 · 14 citations
- Don't Stop the Multi-Party! On Generating Synthetic Written Multi-Party Conversations with ConstraintsNicolò Penzo, Marco Guerini, Bruno Lepri, Goran Glavas et al.AAAI 2026 · 3 citations
- Parrot: Enhancing Multi-Turn Instruction Following for Large Language ModelsYuchong Sun, Che Liu, Kun Zhou, Jinwen Huang et al.ACL 2024
- Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals distinct Multi-Turn Behavior in LLMsClara Lachenmaier, Hannah Bultmann, Sina ZarrießACL 2026
- Benchmarking LLM Tool-Use in the WildPeijie Yu, Wei Liu, Yifan Yang, Jinjian Li et al.ICLR 2026 · 20 citations
