TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games
Yuan Yuan, Muyu He, Muhammad Adil Shahid, Ziyang Li, Jiani Huang, Li Zhang
摘要
This paper introduces TURNABOUTLLM , a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. The framework tasks LLMs with identifying contradictions between testimonies and evidences within long narrative contexts, a challenging task due to the large answer space and diverse reasoning types presented by its questions. We evaluate twelve state-of-the-art LLMs on the dataset, hinting at limitations of popular strategies for enhancing deductive reasoning such as extensive thinking and Chain-of-Thought prompting. The results also suggest varying effects of context size, the number of reasoning step and answer space size on model performance. Overall, TURN-ABOUTLLM presents a substantial challenge for LLMs' deductive reasoning abilities in complex, narrative-rich environments. 1 * Equal contribution. 1 Our resources can be found at https://github.com/zharry29/ turnabout_llm . C o n t r a d ic t io n ! E3 Testimonies Evidences Contradiction! Sahwit claimed he saw the woman dead at 1PM, but the autopsy says she died between 4 and 5PM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri 等ICLR 2024 · 被引用 172 次
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 被引用 38 次
- Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language ModelsNisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja 等EMNLP 2024 · 被引用 6 次
相关 Paper
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi 等NeurIPS 2023 · 被引用 145 次
- Learning Theorem Rationale for Improving the Mathematical Reasoning Capability of Large Language ModelsYu Sheng, Linjing Li, Daniel Dajun ZengAAAI 2025 · 被引用 3 次
- Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMsAbinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury 等ACL 2026
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?Zhanke Zhou, Xiao Feng, Zhaocheng Zhu, Jiangchao Yao 等ICML 2025
- DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured ReasoningQian Wu, Zheyao Gao, Longfei Gou, Qi DouACL 2025 · 被引用 2 次
