Pushing the Limits of Rule Reasoning in Transformers through Natural Language Satisfiability
Kyle Richardson, Ashish Sabharwal
Abstract
Investigating the reasoning abilities of transformer models, and discovering new challenging tasks for them, has been a topic of much interest. Recent studies have found these models to be surprisingly strong at performing deductive reasoning over formal logical theories expressed in natural language. A shortcoming of these studies, however, is that they do not take into account that logical theories, when sampled uniformly at random, do not necessarily lead to hard instances. We propose a new methodology for creating challenging algorithmic reasoning datasets that focus on natural language satisfiability (NLSat) problems. 1 The key idea is to draw insights from empirical sampling of hard propositional SAT problems and from complexity-theoretic studies of language. This methodology allows us to distinguish easy from hard instances, and to systematically increase the complexity of existing reasoning benchmarks such as RuleTaker. We find that current transformers, given sufficient training data, are surprisingly robust at solving the resulting NLSat problems of substantially increased difficulty. They also exhibit some degree of scale-invariance-the ability to generalize to problems of larger size and scope. Our results, however, reveal important limitations too: a careful sampling of training data is crucial for building models that generalize to larger problems, and transformer models' limited scale-invariance suggests they are far from learning robust deductive reasoning algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 977644fb-0193-4dd9-9d93-fb2a276d72a0Cited by top-tier papers9
- Complexity-Based Prompting for Multi-step ReasoningYao Fu, Hao Peng, Ashish Sabharwal, Peter Clark et al.ICLR 2023 · 73 citations
- RECKONING: Reasoning through Dynamic Knowledge EncodingZeming Chen, Gail Weiss, Eric Mitchell, Asli Celikyilmaz et al.NeurIPS 2023 · 21 citations
- What Makes Instruction Learning Hard? An Investigation and a New Challenge in a Synthetic EnvironmentMatthew Finlayson, Kyle Richardson, Ashish Sabharwal, Peter ClarkEMNLP 2022 · 8 citations
- Can Transformers Reason in Fragments of Natural Language?Viktor Schlegel, Kamen V. Pavlov, Ian Pratt-HartmannEMNLP 2022 · 4 citations
- Not all quantifiers are equal: Probing Transformer-based language models' understanding of generalised quantifiersTharindu Madusanka, Iqra Zahid, Hao Li, Ian Pratt-Hartmann et al.EMNLP 2023 · 2 citations
Builds on5
- Probing Natural Language Inference Models through Semantic FragmentsKyle Richardson, Hai Hu, Lawrence S. Moss, Ashish SabharwalAAAI 2020 · 152 citations
- Measuring Systematic Generalization in Neural Proof Generation with TransformersNicolas Gontier, Koustuv Sinha, Siva Reddy, Christopher PalNeurIPS 2020 · 69 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- Do Neural Models Learn Systematicity of Monotonicity Inference in Natural Language?Hitomi Yanaka, Koji Mineshima, Daisuke Bekki, Kentaro InuiACL 2020 · 34 citations
- PRover: Proof Generation for Interpretable Reasoning over RulesSwarnadeep Saha, Sayan Ghosh, Shashank Srivastava, Mohit BansalEMNLP 2020 · 3 citations
Related papers
- Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language ModelsTharindu Madusanka, Ian Pratt-Hartmann, Riza Batista-NavarroACL 2024
- Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability ProblemsTharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann et al.ACL 2025 · 1 citation
- RobustLR: A Diagnostic Benchmark for Evaluating Logical Robustness of Deductive ReasonersSoumya Sanyal, Zeyi Liao, Xiang RenEMNLP 2022 · 6 citations
- Can Transformers Reason Logically? A Study in SAT SolvingLeyan Pan, Vijay Ganesh, Jacob D. Abernethy, Chris Esposo et al.ICML 2025
- SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMsYanxiao Zhao, Yaqian Li, Zihao Bo, Rinyoichi Takezoe et al.ACL 2026
