Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test Oracles
Zihao Xu, Junchen Ding, Yiling Lou, Kun Zhang, Dong Gong, Yuekang Li
Abstract
Large Language Models (LLMs) have achieved significant progress in language understanding and reasoning. Evaluating and analyzing their logical reasoning abilities has therefore become essential. However, existing datasets and benchmarks are often limited to overly simplistic, unnatural, or contextually constrained examples. In response to the growing demand, we introduce SMARTYPAT-BENCH, a challenging, naturally expressed, and systematically labeled benchmark derived from real-world high-quality Reddit posts containing subtle logical fallacies. Unlike existing datasets and benchmarks, it provides more detailed annotations of logical fallacies and features more diverse data. To further scale up the study and address the limitations of manual data collection and labeling, such as fallacy-type imbalance and labor-intensive annotation, we introduce SMARTYPAT, an automated framework powered by logic programming-based oracles. SMARTYPAT utilizes Prolog rules to systematically generate logically fallacious statements, which are then refined into fluent natural language sentences by LLMs, ensuring precise fallacy rep- resentation. Extensive evaluation demonstrates that SMARTYPAT produces fallacies comparable in subtlety and quality to human-generated content and significantly outperforms baseline methods. Finally, experiments reveal insights into LLM capabilities, highlighting that while excessive reasoning steps hinder fallacy detection accuracy, structured reasoning enhances fallacy categorization performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e30d39ed-ecd0-42fe-9f52-68d2e3b030fdBuilds on8
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 865 citations
- DeepCrime: mutation testing of deep learning systems based on real faultsNargiz Humbatova, Gunel Jahangirova, Paolo TonellaISSTA 2021 · 114 citations
- Using Logic Programming to Recover C++ Classes and Methods from Compiled ExecutablesEdward J. Schwartz, Cory F. Cohen, Michael Duggan, Jeffrey Gennari et al.CCS 2018 · 53 citations
- Diagnosing the First-Order Logical Reasoning Ability Through LogicNLIJidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao et al.EMNLP 2021 · 21 citations
- FOLIO: Natural Language Reasoning with First-Order LogicSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi et al.EMNLP 2024 · 18 citations
Related papers
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh et al.EMNLP 2025 · 1 citation
- ORACLE: Optimizing Reasoning Abilities of Large Language Models via Constraint-Led Synthetic Data ElicitationZhuojie Yang, Wentao Wan, Keze WangAAAI 2026
- SLR: Automated Synthesis for Scalable Logical ReasoningLukas Helff, Ahmad Omar, Felix Friedrich, Antonia Wüst et al.ACL 2026 · 6 citations
- Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification InferenceThanh Le-Cong, Bach Le, Toby MurrayACL 2025
- LogicTree: Improving Complex Reasoning of LLMs via Instantiated Multi-step Synthetic Logical DataZehao Wang, Lin F. Yang, Jie Wang, Kehan Wang et al.NeurIPS 2025 · 5 citations
