UniT: A Unified Look at Certified Robust Training against Text Adversarial Perturbation
Muchao Ye, Ziyi Yin, Tianrong Zhang, Tianyu Du, Jinghui Chen, Ting Wang, Fenglong Ma
Abstract
Recent years have witnessed a surge of certified robust training pipelines against text adversarial perturbation constructed by synonym substitutions. Given a base model, existing pipelines provide prediction certificates either in the discrete word space or the continuous latent space. However, they are isolated from each other with a structural gap. We observe that existing training frameworks need unification to provide stronger certified robustness. Additionally, they mainly focus on building the certification process but neglect to improve the robustness of the base model. To mitigate the aforementioned limitations, we propose a unified framework named UniT that enables us to train flexibly in either fashion by working in the word embedding space. It can provide a stronger robustness guarantee obtained directly from the word embedding space without extra modules. In addition, we introduce the decoupled regularization (DR) loss to improve the robustness of the base model, which includes two separate robustness regularization terms for the feature extraction and classifier modules. Experimental results on widely used text classification datasets further demonstrate the effectiveness of the designed unified framework and the proposed DR loss for improving the certified robust accuracy. †
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Scalable Verified Training for Provably Robust Image ClassificationSven Gowal, Krishnamurthy Dvijotham, Robert Stanforth, Rudy Bunel et al.ICCV 2019 · 196 citations
- Generating Natural Language Attacks in a Hard Label Black Box SettingRishabh Maheshwary, Saket Maheshwary, Vikram PudiAAAI 2021 · 128 citations
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li et al.EMNLP 2021 · 114 citations
Related papers
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang et al.S&P 2024 · 41 citations
- Certified Robustness to Programmable Transformations in LSTMsYuhao Zhang, Aws Albarghouthi, Loris D'AntoniEMNLP 2021 · 8 citations
- Generative Adversarial Training with Perturbed Token Detection for Model RobustnessJiahao Zhao, Wenji MaoEMNLP 2023 · 3 citations
- Adversarial Training for Improving Model Robustness? Look at Both Prediction and InterpretationHanjie Chen, Yangfeng JiAAAI 2022 · 31 citations
- DRF: Improving Certified Robustness via Distributional Robustness FrameworkZekai Wang, Zhengyu Zhou, Weiwei LiuAAAI 2024 · 7 citations
