SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT Formulas
Anjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh, Huanmi Tan, Zhanke Zhou, Sanmi Koyejo, Ke Wang, Alex Aiken
摘要
We introduce SATBench, a benchmark for evaluating the logical reasoning capabilities of large language models (LLMs) through logical puzzles derived from Boolean satisfiability (SAT) problems. Unlike prior work that focuses on inference rule-based reasoning, which often involves deducing conclusions from a set of premises, our approach leverages the searchbased nature of SAT problems, where the objective is to find a solution that fulfills a specified set of logical constraints. Each instance in SAT-Bench is generated from a SAT formula, then translated into a puzzle using LLMs. The generation process is fully automated and allows for adjustable difficulty by varying the number of clauses. All 2100 puzzles are validated through both LLM-based and solver-based consistency checks, with human validation on a subset. Experimental results show that even the strongest model, o4-mini, achieves only 65.0% accuracy on hard UNSAT problems, close to the random baseline of 50%. Our error analysis reveals systematic failures such as satisfiability bias, context inconsistency, and condition omission, highlighting limitations of current LLMs in search-based logical reasoning. Our code and data are publicly available at https: //github.com/Anjiang-Wei/SATBench
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Recursive Models for Long-Horizon ReasoningChenxiao Yang, Nati Srebro, Zhiyuan LiICML 2026 · 被引用 6 次
- HardcoreLogic: Challenging Large Reasoning Models with Long-tail Logic Puzzle GamesJingcong Liang, Shijun Wan, Xuehai Wu, Yitong Li 等ICLR 2026 · 被引用 6 次
- Evaluating Robustness of Reasoning Models on Parameterized Logical ProblemsNaïm Es-sebbani, Esteban Marquer, Yakoub Salhi, Zied BouraouiICML 2026 · 被引用 1 次
- ReEfBench: Quantifying the Reasoning Efficiency of LLMsZhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu 等ACL 2026 · 被引用 1 次
- Reasoning Structure of Large Language ModelsFrédéric Berdoz, Luca Lanzendörfer, Fabian Farestam, Roger WattenhoferICML 2026
它引用的顶会 Paper16
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 被引用 325 次
- SatLM: Satisfiability-Aided Language Models Using Declarative PromptingXi Ye, Qiaochu Chen, Isil Dillig, Greg DurrettNeurIPS 2023 · 被引用 126 次
相关 Paper
- SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMsYanxiao Zhao, Yaqian Li, Zihao Bo, Rinyoichi Takezoe 等ACL 2026
- LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language ModelsMihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura 等ACL 2024
- Socrates or Smartypants: Testing Logic Reasoning Capabilities of Large Language Models with Logic Programming-Based Test OraclesZihao Xu, Junchen Ding, Yiling Lou, Kun Zhang 等AAAI 2026 · 被引用 1 次
- LogiConBench: Benchmarking Logical Consistencies of LLMsZheng Chen, Chuan Zhou, Fengxiang Cheng, Tin Po Yip 等ICLR 2026
- InductionBench: LLMs Fail in the Simplest Complexity ClassWenyue Hua, Tyler Wong, Fei Sun, Liangming Pan 等ACL 2025
