Evaluating the Rationale Understanding of Critical Reasoning in Logical Reading Comprehension
Akira Kawabata, Saku Sugawara
摘要
To precisely evaluate a language model’s capability for logical reading comprehension, we present a dataset for testing the understanding of the rationale behind critical reasoning. For questions taken from an existing multiple-choice logical reading comprehension dataset, we crowdsource rationale texts that explain why we should select or eliminate answer options, resulting in 3,003 multiple-choice subquestions that are associated with 943 main questions. Experiments on our dataset show that recent large language models (e.g., InstructGPT) struggle to answer the subquestions even if they are able to answer the main questions correctly. We find that the models perform particularly poorly in answering subquestions written for the incorrect options of the main questions, implying that the models have a limited capability for explaining why incorrect alternatives should be eliminated. These results suggest that our dataset encourages further investigation into the critical reasoning ability of language models while focusing on the elimination process of relevant alternatives.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- FOLIO: Natural Language Reasoning with First-Order LogicSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi 等EMNLP 2024 · 被引用 18 次
- Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language ModelsYuyan Chen, Songzhou Yan, Panjun Liu, Yanghua XiaoACL 2024 · 被引用 8 次
- KRISTEVA: Close Reading as a Novel Task for Benchmarking Interpretive ReasoningPeiqi Sui, Juan Diego Rodriguez, Philippe Laban, Dean Murphy 等ACL 2025 · 被引用 6 次
- BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice BenchmarksNishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie 等ACL 2026 · 被引用 1 次
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales 等ICML 2023 · 被引用 970 次
相关 Paper
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 被引用 38 次
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura 等ACL 2024
- WikiWhy: Answering and Explaining Cause-and-Effect QuestionsMatthew Ho, Aditya Sharma, Justin Chang, Michael Saxon 等ICLR 2023 · 被引用 8 次
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 被引用 110 次
- RMath: A Logic Reasoning-Focused Datasets Toward Mathematical Multistep Reasoning TasksZiyi Hu, Jun Liu, Zhongzhi Liu, Yuzhong Liu 等AAAI 2025 · 被引用 4 次
