MuTual: A Dataset for Multi-Turn Dialogue Reasoning
Leyang Cui, Yu Wu, Shujie Liu, Yue Zhang, Ming Zhou
Abstract
Non-task oriented dialogue systems have achieved great success in recent years due to largely accessible conversation data and the development of deep learning techniques. Given a context, current systems are able to yield a relevant and fluent response, but sometimes make logical mistakes because of weak reasoning capabilities. To facilitate the conversation reasoning research, we introduce Mu-Tual, a novel dataset for Multi-Turn dialogue Reasoning, consisting of 8,860 manually annotated dialogues based on Chinese student English listening comprehension exams. Compared to previous benchmarks for non-task oriented dialogue systems, MuTual is much more challenging since it requires a model that can handle various reasoning problems. Empirical results show that state-of-the-art methods only reach 71%, which is far behind the human performance of 94%, indicating that there is ample room for improving reasoning ability. MuTual is available at https://github. com/Nealcly/MuTual . * Contribution during internship at MSRA. M: Ma'am, you forgot your phone. F: Oh, thanks, I couldn't live without this little thing. M: I know what you mean. It is of great significance to you. So did you enjoy your dinner? F: Oh yes, everything was just perfect. It's so hard to take the whole family out to eat, but your restaurant was perfect. Johnny had his own place to play in and I had time to talk with my sisters and their husbands. ✓ (A) M: Thanks for your compliment for the restaurant. ✘ (B) M: I'm sorry that you don't have a good time. ✘ (C) M: Goodbye brother! Love you. ✘ (D) M: Hurry up honey, or we will be late for the dinner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd57b8de-72b8-47c3-9e1f-6bfb80e31a52Cited by top-tier papers42
- AutoMix: Automatically Mixing Language ModelsPranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju et al.NeurIPS 2024 · 145 citations
- Self-playing Adversarial Language Game Enhances LLM ReasoningPengyu Cheng, Tianhao Hu, Han Xu, Zhisong Zhang et al.NeurIPS 2024 · 120 citations
- The Generative AI Paradox: "What It Can Create, It May Not Understand"Peter West, Ximing Lu, Nouha Dziri, Faeze Brahman et al.ICLR 2024 · 116 citations
- Filling the Gap of Utterance-aware and Speaker-aware Representation for Multi-turn DialogueLongxiang Liu, Zhuosheng Zhang, Hai Zhao, Xi Zhou et al.AAAI 2021 · 57 citations
- Natural Language Inference in Context - Investigating Contextual Reasoning over Long TextsHanmeng Liu, Leyang Cui, Jian Liu, Yue ZhangAAAI 2021 · 57 citations
Related papers
- IRRGN: An Implicit Relational Reasoning Graph Network for Multi-turn Response SelectionJingcheng Deng, Hengwei Dai, Xuewei Guo, Yuanchen Ju et al.EMNLP 2022 · 1 citation
- Exploring Auxiliary Reasoning Tasks for Task-oriented Dialog Systems with Meta Cooperative LearningBowen Qin, Min Yang, Lidong Bing, Qingshan Jiang et al.AAAI 2021 · 9 citations
- MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning EvaluationXiaoyuan Li, Keqin Bao, Yubo Ma, Moxin Li et al.ACL 2026 · 12 citations
- NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven ConversationXiaoyang Wang, Chen Li, Jianqiao Zhao, Dong YuAAAI 2021 · 54 citations
- Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean DialoguesEunsu Kim, Junyeong Park, Juhyun Oh, Kiwoong Park et al.ACL 2026 · 3 citations
