Conditional and Modal Reasoning in Large Language Models
Wesley H. Holliday, Matthew Mandelkern, Cedegao Zhang
摘要
The reasoning abilities of large language models (LLMs) are the topic of a growing body of research in AI and cognitive science. In this paper, we probe the extent to which twenty-nine LLMs are able to distinguish logically correct inferences from logically fallacious ones. We focus on inference patterns involving conditionals (e.g., 'If Ann has a queen, then Bob has a jack') and epistemic modals (e.g., 'Ann might have an ace', 'Bob must have a king'). These inferences have been of special interest to logicians, philosophers, and linguists, since they play a central role in the fundamental human ability to reason about distal possibilities. Assessing LLMs on these inferences is thus highly relevant to the question of how much the reasoning abilities of LLMs match those of humans. All the LLMs we tested make some basic mistakes with conditionals or modals, though zero-shot chain-of-thought prompting helps them make fewer mistakes. Even the best performing LLMs make basic errors in modal reasoning, display logically inconsistent judgments across inference patterns involving epistemic modals and conditionals, and give answers about complex conditional inferences that do not match reported human judgments. These results highlight gaps in basic logical reasoning in today's LLMs. 0% 25% 50% 75% 100% Average correct answer frequency Llama 3.1 Instruct 405B GPT-4 Turbo (2024-04-09) Claude 3.5 Sonnet GPT-4 Turbo (1106) GPT-4 (0613) Llama 3.1 Instruct 70B Gemini 1.5 Pro GPT-4 (0314) GPT-4o (2024-05-13) GPT-4o mini Claude 3 Opus
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- When Reasoning Meets Information Aggregation: A Case Study with Sports NarrativesYebowen Hu, Kaiqiang Song, Sangwoo Cho, Xiaoyang Wang 等EMNLP 2024 · 被引用 10 次
- Evaluating Language Models' Evaluations of GamesKatherine M. Collins, Cedegao E. Zhang, Graham Todd, Lance Ying 等ICLR 2026 · 被引用 5 次
- Logical forms complement probability in understanding language model (and human) performanceYixuan Wang, Freda ShiACL 2025 · 被引用 2 次
- Representations of Fact, Fiction and Forecast in Large Language Models: Epistemics and AttitudesMeng Li, Michael Vrazitulis, David SchlangenACL 2025 · 被引用 1 次
- Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language ModelsBumjin Park, Leejinsil Leejinsil, Jaesik ChoiACL 2025
它引用的顶会 Paper12
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi 等NeurIPS 2023 · 被引用 145 次
相关 Paper
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura 等ACL 2024
- Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMsSiyuan Wang, Zhongyu Wei, Yejin Choi, Xiang RenACL 2024
- No Need for Explanations: LLMs can implicitly learn from mistakes in-contextLisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan 等EMNLP 2025
- Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language ModelsNisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja 等EMNLP 2024 · 被引用 6 次
- LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language ModelsYuxuan Wan, Wenxuan Wang, Yiliu Yang, Youliang Yuan 等EMNLP 2024 · 被引用 10 次
