Towards Verified Robustness under Text Deletion Interventions
Johannes Welbl, Po-Sen Huang, Robert Stanforth, Sven Gowal, Krishnamurthy (Dj) Dvijotham, Martin Szummer, Pushmeet Kohli
Abstract
Neural networks are widely used in Natural Language Processing, yet despite their empirical successes, their behaviour is brittle: they are both over-sensitive to small input changes, and under-sensitive to deletions of large fractions of input text. This paper aims to tackle under-sensitivity in the context of natural language inference by ensuring that models do not become more confident in their predictions as arbitrary subsets of words from the input text are deleted. We develop a novel technique for formal verification of this specification for models based on the popular decomposable attention mechanism by employing the efficient yet effective interval bound propagation (IBP) approach. Using this method we can efficiently prove, given a model, whether a particular sample is free from the under-sensitivity problem. We compare different training methods to address under-sensitivity, and compare metrics to measure it. In our experiments on the SNLI and MNLI datasets, we observe that IBP training leads to a significantly improved verified accuracy. On the SNLI test set, we can verify 18.4% of samples, a substantial improvement over only 2.8% using standard training. * Work done during an internship at DeepMind. 1 This specification is discussed in Section 3. Although a conservative choice, we find it is rarely satisfied. Premise: A little boy in a blue shirt holding a toy. Hypothesis: A boy dressed in blue holds a toy. Entailment (86.4%) Premise: A little boy in a blue shirt holding a toy. Hypothesis: A boy dressed in blue holds a toy. Entailment (91.9%) Original Sample Reduced Sample
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b651a89e-701d-4b35-91a0-27fd7cd6ff2dCited by top-tier papers2
- Robustness to Programmable String Transformations via Augmented Abstract TrainingYuhao Zhang, Aws Albarghouthi, Loris D'AntoniICML 2020 · 17 citations
- Revealing Weaknesses in Text Watermarking Through Self-Information Rewrite AttacksYixin Cheng, Hongcheng Guo, Yangming Li, Leonid SigalICML 2025
Builds on1
Related papers
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel et al.ICCV 2019 · 196 citations
- Understanding Certified Training with Interval Bound PropagationYuhao Mao, Mark Niklas Müller, Marc Fischer, Martin T. VechevICLR 2024 · 26 citations
- Robustness Verification for TransformersZhouxing Shi, Huan Zhang, Kai-Wei Chang, Minlie Huang et al.ICLR 2020 · 131 citations
- Certifiably Adversarially Robust Detection of Out-of-Distribution DataJulian Bitterwolf, Alexander Meinke, Matthias HeinNeurIPS 2020 · 91 citations
- Lipschitz-Certifiable Training with a Tight Outer BoundSungyoon Lee, Jaewook Lee, Saerom ParkNeurIPS 2020 · 23 citations
