Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
Nisarg Patel, Mohith Kulkarni, Mihir Parmar, Aashna Budhiraja, Mutsumi Nakamura, Neeraj Varshney, Chitta Baral
摘要
As Large Language Models (LLMs) continue to exhibit remarkable performance in natural language understanding tasks, there is a crucial need to measure their ability for human-like multi-step logical reasoning. Existing logical reasoning evaluation benchmarks often focus primarily on simplistic single-step or multistep reasoning with a limited set of inference rules. Furthermore, the lack of datasets for evaluating non-monotonic reasoning represents a crucial gap since it aligns more closely with human-like reasoning. To address these limitations, we propose Multi-LogiEval, a comprehensive evaluation dataset encompassing multistep logical reasoning with various inference rules and depths. Multi-LogiEval covers three logic types -propositional, first-order, and non-monotonic consisting of more than 30 inference rules and more than 60 of their combinations with various depths. Leveraging this dataset, we conduct evaluations on a range of LLMs including GPT-4, ChatGPT, Gemini-Pro, Yi, Orca, and Mistral, employing a zero-shot chain-of-thought. Experimental results show that there is a significant drop in the performance of LLMs as the reasoning steps/depth increases (average accuracy of ∼ 68% at depth-1 to ∼ 43% at depth-5). We further conduct a thorough investigation of reasoning chains generated by LLMs which reveals several important findings. We believe that Multi-LogiEval facilitates future research for evaluating and enhancing the logical reasoning ability of LLMs 1 . Rule Combination Context and Question PL Rules: MT, DS Propositions: p: Capture shots in golden hours. q: Photo wins awards. r: Focus on rare wildlife. Context: In wildlife photography, Olivia was certain that if she captured shots in the golden hours, her photos would win awards. However, opportunities varied each day. It was evident that she either captured shots during the golden hours or she focused on rare wildlife, or both. Olivia's latest photos did not win any awards. Question: Is it true that she focused on rare wildlife?
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic CorpusTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaNeurIPS 2024 · 被引用 60 次
- Towards Greater Leverage: Scaling Laws for Efficient Mixture-of-Experts Language ModelsChangxin Tian, Kunlong Chen, Jia Liu, Ziqi Liu 等ICLR 2026 · 被引用 45 次
- Semantic-Aware Logical Reasoning via a Semiotic FrameworkYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li 等ACL 2026 · 被引用 27 次
- VerifyBench: Benchmarking Reference-based Reward Systems for Large Language ModelsYuchen Yan, Jin Jiang, Zhenbang Ren, Yijun Li 等ICLR 2026 · 被引用 18 次
- A Implies B: Circuit Analysis in LLMs for Propositional Logical ReasoningGuanzhe Hong, Nishanth Dikkala, Enming Luo, Cyrus Rashtchian 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 被引用 45 次
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 被引用 38 次
相关 Paper
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura 等ACL 2024
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu 等ICML 2026
- Conditional and Modal Reasoning in Large Language ModelsWesley H. Holliday, Matthew Mandelkern, Cedegao ZhangEMNLP 2024 · 被引用 5 次
- MuSLR: Multimodal Symbolic Logical ReasoningJundong Xu, Hao Fei, Yuhui Zhang, Liangming Pan 等NeurIPS 2025 · 被引用 5 次
- Logical forms complement probability in understanding language model (and human) performanceYixuan Wang, Freda ShiACL 2025 · 被引用 2 次
