NatLogAttack: A Framework for Attacking Natural Language Inference Models with Natural Logic
Zi'ou Zheng, Xiaodan Zhu
Abstract
Reasoning has been a central topic in artificial intelligence from the beginning. The recent progress made on distributed representation and neural networks continues to improve the state-of-the-art performance of natural language inference. However, it remains an open question whether the models perform real reasoning to reach their conclusions or rely on spurious correlations. Adversarial attacks have proven to be an important tool to help evaluate the Achilles' heel of the victim models. In this study, we explore the fundamental problem of developing attack models based on logic formalism. We propose NatLogAttack to perform systematic attacks centring around natural logic, a classical logic formalism that is traceable back to Aristotle's syllogism and has been closely developed for natural language inference. The proposed framework renders both label-preserving and label-flipping attacks. We show that compared to the existing attack models, NatLogAttack generates better adversarial examples with fewer visits to the victim models. The victim models are found to be more vulnerable under the label-flipping setting. NatLogAttack provides a tool to probe the existing and future NLI models' capacity from a key viewpoint and we hope more logicbased attacks will be further explored for understanding the desired property of reasoning. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- Semantics-Aware BERT for Language UnderstandingZhuosheng Zhang, Yuwei Wu, Hai Zhao, Zuchao Li et al.AAAI 2020 · 396 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- Masked Language Model ScoringJulian Salazar, Davis Liang, Toan Q. Nguyen, Katrin KirchhoffACL 2020 · 167 citations
Related papers
- LogicPoison: Logical Attacks on Graph Retrieval-Augmented GenerationYilin Xiao, Jin Chen, Qinggang Zhang, Yujing Zhang et al.ACL 2026
- Logicbreaks: A Framework for Understanding Subversion of Rule-based InferenceAnton Xue, Avishree Khare, Rajeev Alur, Surbhi Goel et al.ICLR 2025
- Mitigating Adversarial Norm Training with Moral AxiomsTaylor Olson, Kenneth D. ForbusAAAI 2023 · 7 citations
- Transitive self-consistency evaluation of NLI models without gold labelsWei Wu, Mark LastEMNLP 2025
- LogiGAN: Learning Logical Reasoning via Adversarial Pre-trainingXinyu Pi, Wanjun Zhong, Yan Gao, Nan Duan et al.NeurIPS 2022 · 19 citations
