Reasoning Gets Harder for LLMs Inside A Dialogue
Ivan Kartác, Mateusz Lango, Ondrej Dusek
Abstract
Large Language Models (LLMs) achieve strong performance on many reasoning benchmarks, yet these evaluations typically focus on isolated tasks that differ from real-world usage in task-oriented dialogue (TOD). In this setting, LLMs must perform reasoning inherently while generating text and adhering to instructions on role, format, and style. This mismatch raises concerns about whether benchmark performance accurately reflects models' reasoning robustness in TOD setting. We investigate how framing reasoning tasks within TOD affects LLM performance by introducing BOUL-DER, a new dynamic benchmark covering eight travel-related tasks that require arithmetic, spatial, and temporal reasoning with both commonsense and formal aspects. Each problem is presented in both isolated and dialogue-based variants, enabling controlled comparison while mitigating data contamination. Experiments on eight LLMs reveal a substantial and consistent performance gap between isolated and dialogue settings. Through ablations and qualitative analysis, we show that this gap is largely driven by the multi-turn nature of dialogue, with additional effects from role conditioning and tool-use requirements. Our results highlight the need to evaluate LLM reasoning in realistic interactive scenarios. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 43ccdc2b-cc9b-4dc3-906f-e0ef78aed4bdBuilds on4
- MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn DialoguesGe Bai, Jie Liu, Xingyuan Bu, Yancheng He et al.ACL 2024 · 35 citations
- SQLWOZ: A Realistic Task-Oriented Dialogue Dataset with SQL-Based Dialogue State Representation for Complex User RequirementsHeng-Da Xu, Xian-Ling Mao, Fanshu Sun, Tian-Yi Che et al.EMNLP 2025
- Rethinking Task-Oriented Dialogue Systems: From Complex Modularity to Zero-Shot Autonomous AgentHeng-Da Xu, Xian-Ling Mao, Puhai Yang, Fanshu Sun et al.ACL 2024
- Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language ModelEmre Can Acikgoz, Jeremiah Greer, Akul Datta, Ze Yang et al.ACL 2025
Related papers
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?Zhanke Zhou, Xiao Feng, Zhaocheng Zhu, Jiangchao Yao et al.ICML 2025
- ACPBench: Reasoning About Action, Change, and PlanningHarsha Kokel, Michael Katz, Kavitha Srinivas, Shirin SohrabiAAAI 2025 · 35 citations
- Do Language Models Have a Common Sense regarding Time? Revisiting Temporal Commonsense Reasoning in the Era of Large Language ModelsRaghav Jain, Daivik Sojitra, Arkadeep Acharya, Sriparna Saha et al.EMNLP 2023 · 17 citations
- BIG-Bench Extra HardMehran Kazemi, Bahare Fatemi, Hritik Bansal, John Palowitch et al.ACL 2025
- MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li et al.ACL 2026 · 12 citations
