LogiConBench: Benchmarking Logical Consistencies of LLMs
Zheng Chen, Chuan Zhou, Fengxiang Cheng, Tin Po Yip, Fenrong Liu, Yisen Wang, Jiajun Chai, Xiaohan Wang, Guojun Yin, Wei Lin, Bo Li, Haoxuan Li, Zhouchen Lin
Abstract
Logical consistency, the requirement that statements remain non-contradictory under logical rules, is fundamental for trustworthy reasoning, yet current LLMs often fail to maintain it even on simple inference tasks. Existing benchmarks for LLM logical consistency are not scalable, not diverse, and not challenging, with state-of-the-art models already surpassing 95% accuracy. LogiConBench is the first benchmark that (1) generates unlimited logical rule combinations with precise labels, (2) provides controllable-depth graphs with explicit reasoning paths, and (3) remains challenging for state-of-the-art LLMs. To achieve this, LogiConBench automatically generates logical graphs where nodes represent symbolic propositions and edges denote reasoning relations. From these graphs, it samples lists of propositions, extracts reasoning paths, determines all consistent label lists, and translates them into diverse natural language expressions. While we release a 280K-sample corpus in this work, the framework can be scaled to generate unlimited data. To strengthen its evaluative significance, we evaluate 14 frontier LLMs on three tasks with varying difficulty levels, and find that the Enumerative task remains extremely challenging, with the best exact accuracy as only 34%. Our code and data are available at https://github.com/Bellafc/LogiConBench.git.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1a56f88c-dc5a-4f40-a7a6-56ee95abc029Cited by top-tier papers1
Ask how each one uses itBuilds on17
- Maieutic Prompting: Logically Consistent Reasoning with Recursive ExplanationsJaehun Jung, Lianhui Qin, Sean Welleck, Faeze Brahman et al.EMNLP 2022 · 72 citations
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 45 citations
- Enhancing Self-Consistency and Performance of Pre-Trained Language Models through Natural Language InferenceEric Mitchell, Joseph J. Noh, Siyan Li, William S. Armstrong et al.EMNLP 2022 · 18 citations
- FOLIO: Natural Language Reasoning with First-Order LogicSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi et al.EMNLP 2024 · 18 citations
- Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve FrameworkJundong Xu, Hao Fei, Meng Luo, Qian Liu et al.ACL 2025 · 11 citations
Related papers
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh et al.EMNLP 2025 · 1 citation
- A²RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark GenerationQingchuan Ma, Yuexiao Ma, Yongkang Xie, Tianyu Xie et al.ICML 2026 · 1 citation
- SLR: Automated Synthesis for Scalable Logical ReasoningLukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst et al.ACL 2026 · 6 citations
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsMohammad Ramezanali, Mo Vazifeh, Paolo SantiEMNLP 2025
- InsLogicBench: An Argumentation Logic Grounded Benchmark for Complex Insurance Claims AdjudicationJin Liu, Yunpeng Liu, Keyi Wang, Jie Shi et al.ACL 2026
