Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
Tharindu Madusanka, Ian Pratt-Hartmann, Riza Batista-Navarro
Abstract
Efforts to apply transformer-based language models (TLMs) to the problem of reasoning in natural language have enjoyed ever-increasing success in recent years. The most fundamental task in this area to which nearly all others can be reduced is that of determining satisfiability. However, from a logical point of view, satisfiability problems vary along various dimensions, which may affect TLMs' ability to learn how to solve them. The problem instances of satisfiability in natural language can belong to different computational complexity classes depending on the language fragment in which they are expressed. Although prior research has explored the problem of natural language satisfiability, the above-mentioned point has not been discussed adequately. Hence, we investigate how problem instances from varying computational complexity classes and having different grammatical constructs impact TLMs' ability to learn rules of inference. Furthermore, to faithfully evaluate TLMs, we conduct an empirical study to explore the distribution of satisfiability problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6db57d2c-98e7-470a-a443-eb5e4172d120Cited by top-tier papers3
- SATBench: Benchmarking LLMs' Logical Reasoning via Automated Puzzle Generation from SAT FormulasAnjiang Wei, Yuheng Wu, Yingjia Wan, Tarun Suresh et al.EMNLP 2025 · 1 citation
- Unravelling the Logic: Investigating the Generalisation of Transformers in Numerical Satisfiability ProblemsTharindu Madusanka, Marco Valentino, Iqra Zahid, Ian Pratt-Hartmann et al.ACL 2025 · 1 citation
- ZebraLogic: On the Scaling Limits of LLMs for Logical ReasoningBill Yuchen Lin, Ronan Le Bras, Kyle Richardson, Ashish Sabharwal et al.ICML 2025
Builds on12
- DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding SharingPengcheng He, Jianfeng Gao, Weizhu ChenICLR 2023 · 394 citations
- What Can Neural Networks Reason About?Keyulu Xu, Jingling Li, Mozhi Zhang, Simon S. Du et al.ICLR 2020 · 281 citations
- Probing Natural Language Inference Models through Semantic FragmentsKyle Richardson, Hai Hu, Lawrence S. Moss, Ashish SabharwalAAAI 2020 · 152 citations
- Differentiable Reasoning on Large Knowledge Bases and Natural LanguagePasquale Minervini, Matko Bosnjak, Tim Rocktäschel, Sebastian Riedel et al.AAAI 2020 · 94 citations
- Differentiable Product Quantization for End-to-End Embedding CompressionTing Chen, Lala Li, Yizhou SunICML 2020 · 81 citations
Related papers
- Can Transformers Reason in Fragments of Natural Language?Viktor Schlegel, Kamen V. Pavlov, Ian Pratt-HartmannEMNLP 2022 · 4 citations
- Pushing the Limits of Rule Reasoning in Transformers through Natural Language SatisfiabilityKyle Richardson, Ashish SabharwalAAAI 2022 · 29 citations
- Transformer Encoder Satisfiability: Complexity and Impact on Formal ReasoningMarco Sälzer, Eric Alsmann, Martin LangeICLR 2025
- Not all quantifiers are equal: Probing Transformer-based language models' understanding of generalised quantifiersTharindu Madusanka, Iqra Zahid, Hao Li, Ian Pratt-Hartmann et al.EMNLP 2023 · 2 citations
- Measuring Systematic Generalization in Neural Proof Generation with TransformersNicolas Gontier, Koustuv Sinha, Siva Reddy, Christopher PalNeurIPS 2020 · 69 citations
