Randomized Smoothing with Masked Inference for Adversarially Robust Text Classifications
Han Cheol Moon, Shafiq R. Joty, Ruochen Zhao, Megh Thakkar, Chi Xu
Abstract
Large-scale pre-trained language models have shown outstanding performance in a variety of NLP tasks. However, they are also known to be significantly brittle against specifically crafted adversarial examples, leading to increasing interest in probing the adversarial robustness of NLP systems. We introduce RSMI, a novel two-stage framework that combines randomized smoothing (RS) with masked inference (MI) to improve the adversarial robustness of NLP systems. RS transforms a classifier into a smoothed classifier to obtain robust representations, whereas MI forces a model to exploit the surrounding context of a masked token in an input sequence. RSMI improves adversarial robustness by 2 to 3 times over existing state-of-the-art methods on benchmark datasets. We also perform in-depth qualitative analysis to validate the effectiveness of the different stages of RSMI and probe the impact of its components through extensive ablations. By empirically proving the stability of RSMI, we put it forward as a practical method to robustly train large-scale NLP models. Our code and datasets are available at https://github.com/Han8931/rsmi_nlp .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- RS-Del: Edit Distance Robustness Certificates for Sequence Classifiers via Randomized DeletionZhuoqun Huang, Neil G. Marchant, Keane Lucas, Lujo Bauer et al.NeurIPS 2023 · 24 citations
- DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative DenoisingZhenhao Li, Huichi Zhou, Marek Rei, Lucia SpeciaACL 2025
- Disentangled Information Bottleneck for Adversarial Text DefenseYidan Xu, Xinghao Yang, Wei Liu, Bao-di Liu et al.EMNLP 2025
Builds on8
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- Towards Robustness Against Natural Language Word SubstitutionsXinshuai Dong, Anh Tuan Luu, Rongrong Ji, Hong LiuICLR 2021 · 63 citations
Related papers
- Model-tuning Via Prompts Makes NLP Models Adversarially RobustMrigank Raman, Pratyush Maini, J. Zico Kolter, Zachary C. Lipton et al.EMNLP 2023 · 7 citations
- Word Level Robustness Enhancement: Fight Perturbation with PerturbationPei Huang, Yuting Yang, Fuqi Jia, Minghao Liu et al.AAAI 2022 · 14 citations
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang et al.S&P 2024 · 41 citations
- Improving Zero-Shot Cross-Lingual Transfer Learning via Robust TrainingKuan-Hao Huang, Wasi Uddin Ahmad, Nanyun Peng, Kai-Wei ChangEMNLP 2021 · 29 citations
- Precisely the Point: Adversarial Augmentations for Faithful and Informative Text GenerationWenhao Wu, Wei Li, Jiachen Liu, Xinyan Xiao et al.EMNLP 2022 · 4 citations
