RobustLR: A Diagnostic Benchmark for Evaluating Logical Robustness of Deductive Reasoners
Soumya Sanyal, Zeyi Liao, Xiang Ren
Abstract
Transformers have been shown to be able to perform deductive reasoning on inputs containing rules and statements written in English natural language. However, it is unclear if these models indeed follow rigorous logical reasoning to arrive at the prediction, or rely on spurious correlation patterns in making decision. A strong deductive reasoning model should consistently understand the semantics of different logical operators. To this end, we present ROBUSTLR, a deductive reasoning-based diagnostic benchmark that evaluates the robustness of language models to minimal logical edits in the inputs and different logical equivalence conditions. In our experiments with RoBERTa, T5, and GPT3, we show that the models trained on deductive reasoning datasets with various logical operations do not perform consistently on the RO-BUSTLR test set, thus showing that the models are not robust to our proposed logical perturbations. Further, we observe that the models find it especially hard to learn logical negation operator. Our results demonstrate the shortcomings of current language models in logical reasoning, and call for the development of better inductive biases to teach the logical semantics to language models. All the datasets and code base have been made publicly available. 1 f1: Charlie is tall. r1: Erin is kind, if Charlie is tall. statement: Erin is kind. Label: True f1: Charlie is tall. r1: Erin is kind, if Charlie is tall and round. statement: Erin is kind. Label: Unknown (a) Original Theory (b) Conjunction Perturbation f1: Charlie is tall. r1: Erin is kind, if Charlie is tall or round.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi et al.NeurIPS 2023 · 145 citations
- RECKONING: Reasoning through Dynamic Knowledge EncodingZeming Chen, Gail Weiss, Eric Mitchell, Asli Celikyilmaz et al.NeurIPS 2023 · 21 citations
- LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language ModelsKang He, Kaushik RoyEMNLP 2025
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- ReClor: A Reading Comprehension Dataset Requiring Logical ReasoningWeihao Yu, Zihang Jiang, Yanfei Dong, Jiashi FengICLR 2020 · 325 citations
- MuTual: A Dataset for Multi-Turn Dialogue ReasoningLeyang Cui, Yu Wu, Shujie Liu, Yue Zhang et al.ACL 2020 · 115 citations
- Measuring Systematic Generalization in Neural Proof Generation with TransformersNicolas Gontier, Koustuv Sinha, Siva Reddy, Christopher PalNeurIPS 2020 · 69 citations
Related papers
- Diagnosing the First-Order Logical Reasoning Ability Through LogicNLIJidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao et al.EMNLP 2021 · 21 citations
- FaiRR: Faithful and Robust Deductive Reasoning over Natural LanguageSoumya Sanyal, Harman Singh, Xiang RenACL 2022 · 49 citations
- Pushing the Limits of Rule Reasoning in Transformers through Natural Language SatisfiabilityKyle Richardson, Ashish SabharwalAAAI 2022 · 29 citations
- MME-Reasoning: A Broad-Spectrum Benchmark for Evaluating Logical Reasoning in MLLMsJiakang Yuan, Tianshuo Peng, Yilei Jiang, Yiting Lu et al.ICML 2026
- Learning Deductive Reasoning from Synthetic Corpus based on Formal LogicTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaICML 2023 · 45 citations
