Logical Reasoning with Span-Level Predictions for Interpretable and Robust NLI Models
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Marek Rei
Abstract
Current Natural Language Inference (NLI) models achieve impressive results, sometimes outperforming humans when evaluating on in-distribution test sets. However, as these models are known to learn from annotation artefacts and dataset biases, it is unclear to what extent the models are learning the task of NLI instead of learning from shallow heuristics in their training data.We address this issue by introducing a logical reasoning framework for NLI, creating highly transparent model decisions that are based on logical rules. Unlike prior work, we show that improved interpretability can be achieved without decreasing the predictive accuracy. We almost fully retain performance on SNLI, while also identifying the exact hypothesis spans that are responsible for each model prediction.Using the e-SNLI human explanations, we verify that our model makes sensible decisions at a span level, despite not using any span labels during training. We can further improve model performance and the span-level decisions by using the e-SNLI explanations during training. Finally, our model is more robust in a reduced data setting. When training with only 1,000 examples, out-of-distribution performance improves on the MNLI matched and mismatched validation sets by 13% and 16% relative to the baseline. Training with fewer observations yields further improvements, both in-distribution and out-of-distribution.<br/>
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f51dbe57-d8a2-43f5-ba26-42bacbce0038Cited by top-tier papers4
- QA-NatVer: Question Answering for Natural Logic-based Fact VerificationRami Aly, Marek Strong, Andreas VlachosEMNLP 2023 · 7 citations
- Extractive Fact Decomposition for Interpretable Natural Language Inference in one Forward PassNicholas Popovic, Michael FärberEMNLP 2025 · 1 citation
- InfoLossQA: Characterizing and Recovering Information Loss in Text SimplificationJan Trienes, Sebastian Joseph, Jörg Schlötterer, Christin Seifert et al.ACL 2024
- Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More MoreArvid FrydenlundACL 2025
Builds on9
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 136 citations
- Leap-Of-Thought: Teaching Pre-Trained Models to Systematically Reason Over Implicit KnowledgeAlon Talmor, Oyvind Tafjord, Peter Clark, Yoav Goldberg et al.NeurIPS 2020 · 119 citations
- Supervising Model Attention with Human Explanations for Robust Natural Language InferenceJoe Stacey, Yonatan Belinkov, Marek ReiAAAI 2022 · 52 citations
Related papers
- LIREx: Augmenting Language Inference with Relevant ExplanationsXinyan Zhao, V. G. Vinod VydiswaranAAAI 2021 · 41 citations
- Weakly Supervised Explainable Phrasal Reasoning with Neural Fuzzy LogicZijun Wu, Zi Xuan Zhang, Atharva Naik, Zhijian Mei et al.ICLR 2023 · 5 citations
- Rule Discovery for Natural Language Inference Data Generation Using Out-of-Distribution DetectionJuyoung Han, Hyunsun Hwang, Changki LeeEMNLP 2025
- Improving the robustness of NLI models with minimax trainingMichalis Korakakis, Andreas VlachosACL 2023 · 4 citations
- FLamE: Few-shot Learning from Natural Language ExplanationsYangqiaoyu Zhou, Yiming Zhang, Chenhao TanACL 2023 · 8 citations
