Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
Xinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang, Binghui Wang, Zhongjie Ba, Kui Ren
Abstract
The language models, especially the basic text classification models, have been shown to be susceptible to textual adversarial attacks such as synonym substitution and word insertion attacks. To defend against such attacks, a growing body of research has been devoted to improving the model’s robustness. However, providing provable robustness guarantees instead of empirical robustness is still widely unexplored. In this paper, we propose Text-CRS, a generalized certified robustness framework for natural language processing (NLP) based on randomized smoothing. To our best knowledge, existing certified schemes for NLP can only certify the robustness against ℓ0 perturbations in synonym substitution attacks. Representing each word-level adversarial operation (i.e., synonym substitution, word reordering, insertion, and deletion) as a combination of permutation and embedding transformation, we propose novel smoothing theorems to derive robustness bounds in both permutation and embedding space against such adversarial operations. To further improve certified accuracy and radius, we consider the numerical relationships between discrete words and select proper noise distributions for the randomized smoothing. Finally, we conduct substantial experiments on multiple language models and datasets. Text-CRS can address all four different word-level adversarial operations and achieve a significant accuracy improvement. We also provide the first benchmark on certified accuracy and radius of four word-level operations, besides outperforming the state-of-the-art certification against synonym substitution attacks.1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ddcabff-88ea-40ae-944e-f656fc2cef2aCited by top-tier papers18
- Exposing the Deception: Uncovering More Forgery Clues for Deepfake DetectionZhongjie Ba, Qingyu Liu, Zhenguang Liu, Shuang Wu et al.AAAI 2024 · 101 citations
- Fight Back Against Jailbreaking via Prompt Adversarial TuningYichuan Mo, Yuji Wang, Zeming Wei, Yisen WangNeurIPS 2024 · 90 citations
- PromptCARE: Prompt Copyright Protection by Watermark Injection and VerificationHongwei Yao, Jian Lou, Zhan Qin, Kui RenS&P 2024 · 43 citations
- Attention! Your Vision Language Model Could Be Maliciously ManipulatedXiaosen Wang, Shaokang Wang, Zhijin Ge, Yuyang Luo et al.NeurIPS 2025 · 13 citations
- Distributed Backdoor Attacks on Federated Graph Learning and Certified DefensesYuxin Yang, Qiang Li, Jinyuan Jia, Yuan Hong et al.CCS 2024 · 8 citations
Builds on28
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu et al.ACL 2020 · 188 citations
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 128 citations
Related papers
- Certified Robustness to Programmable Transformations in LSTMsYuhao Zhang, Aws Albarghouthi, Loris D'AntoniEMNLP 2021 · 8 citations
- GSmooth: Certified Robustness against Semantic Transformations via Generalized Randomized SmoothingZhongkai Hao, Chengyang Ying, Yinpeng Dong, Hang Su et al.ICML 2022 · 27 citations
- Certifying Language Model Robustness with Fuzzed Randomized Smoothing: An Efficient Defense Against Backdoor AttacksBowei He, Lihao Yin, Hui-Ling Zhen, Jianping Zhang et al.ICLR 2025
- TSS: Transformation-Specific Smoothing for Robustness CertificationLinyi Li, Maurice Weber, Xiaojun Xu, Luka Rimanic et al.CCS 2021 · 21 citations
- Certified Robustness Against Natural Language Attacks by Causal InterventionHaiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu et al.ICML 2022 · 43 citations
