TurnaboutLLM: A Deductive Reasoning Benchmark from Detective Games
Yuan Yuan, Muyu He, Muhammad Adil Shahid, Ziyang Li, Jiani Huang, Li Zhang
Abstract
This paper introduces TURNABOUTLLM , a novel framework and dataset for evaluating the deductive reasoning abilities of Large Language Models (LLMs) by leveraging the interactive gameplay of detective games Ace Attorney and Danganronpa. The framework tasks LLMs with identifying contradictions between testimonies and evidences within long narrative contexts, a challenging task due to the large answer space and diverse reasoning types presented by its questions. We evaluate twelve state-of-the-art LLMs on the dataset, hinting at limitations of popular strategies for enhancing deductive reasoning such as extensive thinking and Chain-of-Thought prompting. The results also suggest varying effects of context size, the number of reasoning step and answer space size on model performance. Overall, TURN-ABOUTLLM presents a substantial challenge for LLMs' deductive reasoning abilities in complex, narrative-rich environments. 1 * Equal contribution. 1 Our resources can be found at https://github.com/zharry29/ turnabout_llm . C o n t r a d ic t io n ! E3 Testimonies Evidences Contradiction! Sahwit claimed he saw the woman dead at 1PM, but the autopsy says she died between 4 and 5PM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 325 citations
- MuSR: Testing the Limits of Chain-of-thought with Multistep Soft ReasoningZayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri et al.ICLR 2024 · 172 citations
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
- Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language ModelsNisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja et al.EMNLP 2024 · 6 citations
Related papers
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi et al.NeurIPS 2023 · 145 citations
- Learning Theorem Rationale for Improving the Mathematical Reasoning Capability of Large Language ModelsYu Sheng, Linjing Li, Daniel Dajun ZengAAAI 2025 · 3 citations
- Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMsAbinitha Gourabathina, Inkit Padhi, Manish Nagireddy, Subhajit Chaudhury et al.ACL 2026
- From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?Zhanke Zhou, Xiao Feng, Zhaocheng Zhu, Jiangchao Yao et al.ICML 2025
- DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured ReasoningQian Wu, Zheyao Gao, Longfei Gou, Qi DouACL 2025 · 2 citations
